In 2009, then-high school senior Zohran Mamdani selected the boxes “Asian” and “African-American” to define his race on his Columbia University application. Sixteen years later, however, much of his 2025 New York City mayoral campaign hinged on his Muslim and South Asian ancestry.
Mamdani, whose application leaked after an internal Columbia data breach in July 2025, stated that while he does not consider himself African-American, he stands by his choice. Born to Indian parents in Uganda, Mamdani holds dual Ugandan-American citizenship.
“Most college applications don’t have a box for Indian-Ugandans, so I checked multiple boxes trying to capture the fullness of my background,” Mamdani said.
While some argue he took advantage of affirmative-action policies, the mayor of New York City’s struggle to conceptualize his complex cultural identity is not unique. When Americans are asked to define themselves, the answers span thousands of distinct identities and combinations.
The Office of Budget and Management (OMB), however, condenses those responses into five categories on the U.S. Census. Every 10 years, the Census collects data, attempting to paint an accurate portrait of America. These standardized racial categories are used to allocate resources and track inequality. In the process, racial data aggregation — statistical invisibility or the practice of breaking down broad racial categories into specific ethnic groups — has the potential to obscure the very disparities these datasets aim to illuminate.
Today, about half of Americans believe the Census categories reflect their identity “very well.” Those same citizens have struggled to identify their race and ethnicity with shifting social norms and eras in American history. Between the 2010 and 2020 Censuses, 9.8 million Americans switched their race or identity.
“If social science evidence is correct, people are constantly experiencing and negotiating their racial and ethnic identities in interactions with people and institutions, and in personal, local, national, and historical context,” the University of Michigan study said.
The nation’s interpretation of race and ethnic categories has fluctuated since the 1790 conception of the decennial Census, the nation’s principal registrar of population data. Initially, the Census reflected the social stratification of the 18th century, perpetuating the dehumanization of Black people with the Three-Fifths Compromise. The legislation diminished Black people to three-fifths of a person to boost Southern states’ population figures and grant them more representation in the House of Representatives and Electoral College. The Census classified Americans into two clear-cut distinctions: free or slave. Congress expanded the Census to include “free colored persons” as a category in 1820, followed by the introduction of Chinese and Japanese categories in the late 19th century.
In 1977, the OMB added four racial categories: “American Indian or Alaskan Native,” “Black,” “White,” and “Asian or Pacific Islander,” which was later split into “Asian” and “Native Hawaiian or Pacific Islander” in 1997, establishing the five primary racial groups used in the Census today.
Katie Weng, a data science major at Carnegie Mellon University, said that the Census categories don’t accurately represent the population.
“They’re way too broad and don’t match how many people actually think about their identity,” Weng said. “Categories like ‘Asian’ or ‘Black’ cover incredibly diverse groups with very different histories and outcomes, but that gets flattened in the data.”
While it was not initially its own racial category, the OMB added “Hispanic” as an ethnicity to the 1980 Census after pressure from advocacy groups such as the Mexican American Legal Defense Fund. Because the term “Hispanic” describes people with Spanish-speaking heritage rather than skin color or physical appearance, the OMB classified the category as an ethnicity rather than a race, hinging on the idea that a person of any race can identify as Hispanic.
In the 2020 Census, 28 million Americans selected the single category “some other race,” essentially making their racial identities and specific equity needs invisible to the federal government. The states with the highest concentration of those who selected “some other race” were along the Southern border, a region with a large population of Hispanic individuals. Because 67% of Latino Americans perceive being Latino as part of their racial background, some researchers believe that Latinos make up the largest proportion of “some other race” individuals, though their data remains unknown.
To rectify this issue, the new revisions to the OMB’s 1997 Statistical Policy Directive No. 15 (SPD-15) will consider “Hispanic” a race in the 2030 Census and onwards. However, some oppose the addition of “Hispanic” as a race. The conflation of Latino ethnicity with race may precipitate inaccurate self-identification for many, causing an undercount for the community. For example, some Afro-Latinx might select Latino as their race and not mark Black separately, as they would if Latino were an ethnicity option.
Some scholars also advocate for the adoption of a “street race” question on the Census, referring to one’s perceived race based on physical appearance. Its inclusion could provide data on discrimination and serve as a tool for civil rights research.
Ana Ortizoga, a health equity advisor at the Pan American Health Organization, stressed the importance of disaggregating racial and ethnic categories.
“Many disparities in health, usually ascribed to socioeconomic status, sometimes mask deeper inequalities related to racisms or other forms of discrimination,” Ortizoga said. “These inequities can be uncovered when you further disaggregate data by categories of race and ethnicity.”
While changes to how race is measured could strengthen the data itself, existing statistics reveal significant disparities within racial categories. Asian Americans, while often seen as a “model minority,” are labeled with a stereotype that all Asians are affluent and well-educated. In reality, Asian Americans demonstrate a bimodal distribution of wealth in the top and bottom 10% of the nation’s income spectrum. Mongolian and Burmese groups, specifically, have a poverty rate higher than twice the national average.
The all-encompassing term “Asian” can also merge ethnic groups with varying health challenges. In 2022, the transnational organization AF3IRM launched “Kanlungan,” meaning refuge in Tagalog, as a digital sanctuary. AF3IRM pioneered “Kanlungan” in response to federal data that reported a lower mortality rate for Asian Americans during the pandemic, which did not reflect the high proportion of Filipino nurses among COVID nurse deaths. The data wall memorializes Filipino healthcare workers who lost their lives to the COVID-19 virus.
When the federal government reported that Asian Americans had comparatively low COVID-19 mortality rates, the numbers suggested resilience. In reality, Filipino nurses, who make up 4% of nurses — the country’s largest nursing union — were dying at alarming rates, a disparity buried under a single racial category.
Similarly, the “Black” category maintains discrepancies, largely because the Census doesn’t differentiate between African Americans, African immigrants and Afro-Caribbeans, groups with differing ethnic, health-related and linguistic characteristics. In several metropolitan cities like Miami, Boston and Atlanta, African immigrants have a lower poverty rate than African Americans, though widespread data across the U.S. is not available due to aggregation.
In the U.S. Census, American Indian and Alaska Native (AI/AN) populations appear as a single demographic, limiting the cultural nuance and sovereignty of distinct nations. Misracialization is another risk of Census data collection contributing to aggregated data — a practice that identifies AI/AN groups as other racial or ethnic identities.
AI/AN individuals with multiracial identity or Hispanic heritage are grouped into broad categories like “Hispanic/Latino” or “Two or More Races” rather than Indigenous. The undercount limits understanding of AI/AN trends, a community that demonstrates 1.7 times the rate of suicides and 6.5 times the alcohol related deaths compared to non-Hispanic white people. Even among those classified as AI/AN, there are differences regarding tribal affiliation.
The U.S. has 574 federally recognized tribes, not including state-identified and unregistered tribes. Although former president Bill Clinton issued an executive order to coordinate policies, including Census distribution with recognized U.S. tribes, “data genocide” — the erasure of AI/AN groups in federal data pools — remains an ongoing concern within the community.
According to data collected from reservations where obstetrical care — medical support for pregnancy, delivery and the postpartum period — is sparse, AI/AN women have the second-highest maternal mortality rate in the U.S. after Black women. The accuracy of these statistics is unclear due to limited epidemiological data from regional centers and the Census, hindering national efforts to improve maternal health for Indigenous women.
Weng said that disaggregating data is often considered optional by many institutions — perpetuating a pattern of negligence that can have tangible impacts on minority communities.
“Policymakers usually work with whatever data already exists, even if it’s messy, oversimplified or missing important context,” Weng said. “While fairness gets discussed in theory, the actual decisions end up being based on categories and metrics that don’t reflect the nuance everyone claims to care about.”
Historically, Middle Eastern and North African (MENA) individuals have been excluded from Census data collections, forced to check “White,” “Black” or “Asian” boxes. However, they may not identify with the provided categories. When the first wave of Arab immigrants arrived in the U.S. in the early 1800s, they fought to be classified as white to bypass immigration restrictions. The federal government classified MENA people as white in 1977, setting a legal precedent for today’s data collection.
Junior Rosha Rizi, who is Iranian-American, believes the “white” category fails to reflect her heritage and the experiences that shape her daily life. In elementary school, she recalled being confused when teachers told her to put down white when filling out forms.
“I’m not physically white,” Rizi said. “That’s the reality — it’s just ignoring the fact that it’s a region with a completely different culture and identity.”
The aggregation of MENA datasets has the potential to impact equity. While there is extensive information on achievement disparities between white and Black students, little is known regarding MENA students due to the lack of available data. Anti-Arab and Muslim discrimination is also harder to prove statistically because white people show low rates of discrimination in aggregate data.
Rizi said that the exclusion of a MENA box prevents members of her community from getting the various forms of assistance they need from the government.
“A lot of people who immigrate from the Middle East are coming from war-torn backgrounds,” Rizi said. “My parents came because there was a war in Iran, and they came with very little money. They didn’t have access to the stuff that they needed from the government, like healthcare.”
MENA advocacy groups, such as the Arab American Institute, helped push for the revision of SPD-15, which will implement a new MENA ethnic category for the 2030 Census, no longer requiring those who identify as MENA to classify themselves as “white.” Government agencies will also be required to provide detailed subcategory reports to capture diversity more accurately within broad racial groups.
The revision will redistribute the country’s recorded data pool and increase equity awareness towards MENA communities across various policy areas. Updates to the Census, however, are currently under threat. As of December 2025, the OMB is reviewing the revision parameters. White House Chief Statistician Mark Calabria told NPR that the evaluation process is still in a premature phase, providing the first public confirmation that the administration may not use the latest racial and ethnic category changes for the 2030 Census.
Strides in racial data processing will not be uniform outside of the Census. Non-federal systems, including hospitals, universities and private companies, are not necessarily obligated to report racial discrepancies beyond the OMB’s 1997 SPD-15 policy standards.
Some states and policymakers, however, are working to disaggregate data. A 2024 California law allowed state agency Black applicants to identify themselves in a range of diverse categories, including descendants of enslaved people in the U.S., African Blacks, American Freedmen and Caribbean Blacks. In 2021, New York passed a law requiring every state agency to include at least 20 Asian and six Native Hawaiian and Pacific Islander groups in Asian data collection.
Ortizoga said that there are also risks that accompany disaggregating data.
“Too many categories could make population estimations very small numbers that cannot be analyzed afterwards,” Ortizoga said. “However, an appropriate number of race and ethnicity categories can help to ‘unmask’ health disparities that are rooted in discriminatory practices such as residential segregation and treatment in the health system.”
There are several measures institutions can take to ensure racial analytics are accurately represented. At an individual level, clinicians and researchers can collect and use disaggregated data in intake surveys and studies. When using secondary data, researchers can reference sources such as the National Health Interview Survey or American Community Survey, which provide ethnic subcategories for racial data. On the institutional level, health systems can use algorithms to identify subgroups. Bayesian Improved Surname Coding, a system that predicts an individual’s race and ethnic subcategory based on their surname and residential address, can qualify as a reference for missing data points.
Despite emerging methods to refine racial data processing, the impact of aggregation persists at the individual level. For millions of people like Rizi, census categories will continue to shape how they are reflected in, or omitted from, public data.
“I know that I have this background,” Rizi said. “I have a specific culture, but am I supposed to just erase that part of my identity and check a box where I don’t necessarily fit?”
