Who gets to define ethnicity? Identity and ethnicity data
Categories: Research using linked data, Blogs
28 August 2026
Note: This blog contains terminology relating to race and ethnicity that some readers may find offensive. These terms are included where relevant to accurately discuss language, classifications and social attitudes.
This blog is the second in a series by Maggie, the 2026 Communications & Engagement Intern at ADR UK. Maggie explores how people define their own ethnicity, and what this means when complex identities have to be translated into data. The reflections in this blog are Maggie's own and are intended to encourage thoughtful discussion around the complexities of race, ethnicity and ethnicity data.
Ethnicity is not simply something defined by fixed categories or assigned by institutions. It is also something we understand and define for ourselves. So, what happens when something as personal, complex and sometimes fluid as ethnic identity needs to be recorded as data?
There can be a long journey between how somebody understands their own ethnicity and the ethnicity variable that eventually appears in a dataset. In this blog, I explore why people’s ability to define and report their own ethnicity matters, what happens when those identities are translated into categories, and what researchers need to consider when interpreting the resulting data.
From self-definition to self-reported data
To understand that process, it helps to distinguish between three closely related ideas:
Self-definition is the inherently personal process through which we understand and define our own identity, values and character.
Self-identification is how we describe those aspects to others.
When we provide information about our own identity for it to be recorded, this becomes self-reported data.
This distinction is important because ethnicity categories have not always been defined by the people they describe. As I explored in the previous blog, “ethnicity data does not exist in a vacuum. The categories used in datasets today are shaped by historical, social, and institutional contexts”. Some have been prescribed by institutions without consulting the people being categorised.
Today, self-identification and self-reporting are therefore important principles in the collection of ethnicity data. UK Government guidance recommends collecting self-reported ethnicity wherever possible, while Office for National Statistics (ONS) guidance describes ethnic group as a self-identification measure.
Why self-definition and self-identification matter
Choice and agency
Self-definition gives people greater agency over how they understand their own ethnic identity, rather than feeling their identity is something determined entirely by other people, institutions or stereotypes. Self-identification allows us to express that understanding to others.
Personally, this has been particularly liberating during my time at university. It has helped me develop greater pride, excitement and curiosity about my culture and ethnic origins. Understanding and defining my identity has been an active journey of reconnection, creativity, education and relationships, including exploring archives and learning more about the histories that have shaped who I am.
The writer and scholar, bell hooks, wrote in her book Rock My Soul that “in the contemporary world Black identities are diverse and complex”, and that holistic education is key to developing Black self-esteem and self-love. For me, self-definition has been part of that journey.
When I self-identify as a Black woman, that means much more to me than my skin colour. My identity is shaped by both my race and my ethnicity. While being Black speaks to my racial identity and shared experiences of how I am perceived and treated in society, being Yoruba/Nigerian connects me to a distinct ethnic heritage, history, traditions and community. My identity, or my “essence”, is a collage of the Black music genres I listen to, the way my hair is braided, the books I read, the films I watch, the food I eat, and the relationships and communities around me. These are all ways in which I interact with and understand my culture, beliefs, ancestry and personality, including aspects of identity that legacies of colonialism, misogyny and racism have attempted to erase or misrepresent.
Better reflecting people's identities
Allowing people to report their own ethnicity also matters for the quality of the information that is recorded about them.
Giving somebody the opportunity to report their own ethnicity recognises that they have greater knowledge and authority over their identity than somebody attempting to assign an ethnicity to them from the outside. It can reduce the risk of misclassification based on somebody else's assumptions and allow data to better reflect changing identities and wider social change.
This is not just a theoretical preference. The Race Disparity Unit has found that third-party ethnicity information is generally less reliable for measuring a person’s own ethnic identity than self-reported information, because it is based on someone else’s interpretation. Its guidance therefore recommends maximising opportunities for people to report the ethnicity they identify with, including through manual write-in options or more detailed categories where appropriate.
For researchers, better information about how people identify themselves can help build a richer, long-term picture of the populations they are studying and the inequalities they may experience.
But there is still a journey between defining an identity, describing it to others, and recording it as data. Self-reporting gives people greater say in that process, but recording ethnicity still means translating complex identities into standardised information that can be stored and used.
Categorising identity
Even when we report our own ethnicity, we are usually asked to do so using the categories available on a particular form or within a particular system. When I fill out a form, for example, I might select Black African or Black British. Those are both ways I might identify myself, but neither captures everything my ethnicity and culture mean to me.
That does not necessarily mean the categories are inaccurate. Rather, it illustrates the difference between the richness of an identity and the particular piece of information that can be captured in a standardised data field. A category can meaningfully describe part of somebody's identity without describing all of it.
The categories available to us, and what those categories signify, are also shaped by place and history.
For example, in Northern Ireland, ethnicity data is recorded in a way that complies with the Good Friday Agreement. National identity questions and categorisations are asked in a way that no one must choose between being British, Irish, and Northern Irish. This emphasises how categorising ethnic identity can be complex, political and situational. It also highlights that ethnicity data not just derived from how you define yourself, but also what questions institutions are asking.
Another good example is shown through the South African singer Tyla. She has described herself as “Coloured” – a recognised identity with a particular historical and cultural meaning in South Africa. Her use of the term prompted debate internationally because the same word has different, and often offensive, historical connotations in countries including the UK and United States. In another classification system, she might instead be categorised using a label such as Mixed Ethnicity.
What ultimately appears in a dataset therefore reflects an interaction between how somebody understands themselves, how they choose to describe themselves in a particular context, and the categories or rules of the system collecting that information.
The challenges of using self-reported ethnicity data
Self-reported ethnicity brings data closer to how people understand themselves, but it cannot make complexities disappear.
Balancing detail and comparison
Researchers often need standardised information that can be compared across people, places and time.
Imagine a dataset in which everybody could describe their ethnicity in completely different terms. It might capture people's identities in rich detail, but it could become extremely difficult to identify patterns across a population or compare people's experiences. Disadvantaged groups, who require greater policy attention, may go overlooked.
At the other extreme, if everybody had to fit themselves into only two or three very broad categories, the data would be easier to compare, but important and granular differences between communities could disappear.
Ethnicity data collection therefore involves balancing opportunities for people to describe themselves meaningfully with the need to create information that can be analysed consistently. This is reflected in the ONS's use of harmonised ethnicity questions, which are designed to support consistency and comparability between different sources.
This is particularly important in administrative data research. Administrative data is usually collected while delivering public services rather than specifically for research. Ethnicity information may therefore have been recorded by different organisations, at different times, for different purposes and using different methods and classifications.
Missing or incomplete information
Researchers should also be aware of who may be missing or inadequately represented in ethnicity data. Gaps can arise when individuals choose not to disclose their ethnicity, ethnicity information is not collected or transferred between systems, or the available response options do not reflect how people identify. This can mean that some communities are underrepresented in a dataset or are present but cannot be identified as distinct groups.
If missingness or inadequate representation is concentrated among particular groups or in particular settings, it can affect the conclusions researchers draw from a dataset. Researchers should therefore consider not only which ethnicity categories are represented, but also who may be absent, who may be hidden within broader categories and why.
These limitations can also carry through into linked and synthetic datasets, or into subsequent analysis using AI. Linking, generating or analysing data cannot automatically recover identities or distinctions that were never collected or were obscured in the source data, and may reproduce existing gaps or patterns of underrepresentation.
The examples below, drawing on Indo-Caribbean, Latin American and Hispanic, and Arab and South West Asian and North African communities, illustrate how missing or inadequate representation in ethnicity data can affect what we know about different groups and their experiences.
Indo-Caribbean identities
Indra Nauth, a British-born Indo-Guyanese woman, has written about the difficulty of choosing an ethnicity category that accurately represents her identity. She highlights that it is difficult to know how many Indo-Caribbean people live in the UK because they are not specifically represented within census ethnicity categories.
Her experience illustrates how standard classifications can struggle to capture identities that span different ethnic, cultural and geographic backgrounds. Indo-Caribbean people may therefore be recorded within broader categories without being identifiable as a distinct group, making their particular experiences and needs less visible in the data. This can have implications for research, policymaking and public service provision.
Latin American and Hispanic identities
Hannah Manzur highlights that Latin American and Hispanic people are not represented as a distinct ethnicity category in much UK data collection. Instead, respondents may be captured within broader categories such as “White” or “Other”, potentially masking differences in their experiences and needs.
Research discussed by Manzur found that Latin American and Hispanic respondents reported higher levels of victimisation and fear of violence when analysed as a distinct group. This demonstrates how the categories used to collect and analyse ethnicity data can affect the patterns and inequalities that researchers are able to identify.
Arab & South West Asian and North African identities
Hannah Manzur suggests that the use of the broad category “Arab” in some UK datasets should be reviewed. “Arab” encompasses people with a wide range of national, ethnic and cultural identities, while South West Asia and North Africa (SWANA) is a geographical term covering a range of communities. The terms are therefore not interchangeable: not everyone from the SWANA region identifies as Arab, and some people may find that the ethnic categories available to them do not reflect how they understand their own identity.
Existing categories can also present difficulties for people from North Africa, who may identify as African without identifying as Black. Labels such as “Black African” illustrate how ethnicity classifications can combine racial, geographic and cultural identities in ways that do not always reflect how individuals identify themselves. Distinct communities may consequently be grouped together or become less visible within the data.
This matters for research and policymaking. If Arab and other SWANA communities cannot be identified accurately within datasets, important differences in their experiences and needs may be obscured, making inequalities harder to identify and address.
The same person might be recorded differently
Even when ethnicity is self-reported, the same person is not necessarily recorded in exactly the same way throughout their life.
A simple analogy is my name. My full name is Margaret, but in everyday life I almost always go by Maggie. Perhaps one day I'll decide I prefer Mags. Depending on the context, one dataset might record me as Margaret and another as Maggie. I am still the same person, but the information recorded about me differs. This might even skew data, giving a less reliable representation of how many Maggies, Margarets, and Mags there are in a population.
Ethnicity is of course much more complex than a nickname, but something similar can happen in the data. Somebody may select one ethnicity category at one point in their life and another later. They might identify with more than one category, find that different options are available on different forms, or change how they choose to describe themselves.
Evidence suggests this happens in practice. Race Disparity Unit analysis of the Understanding Society youth survey found that around one in ten children in the study reported a different ethnicity at least once, although the figures were not designed to be representative of the wider population. The report also cites research comparing UK census records from 2001 and 2011, which found that 4% of people who provided ethnicity information reported it differently between the two censuses.
A difference in somebody's recorded ethnicity does not necessarily mean their identity has changed – or that somebody has made a mistake. It may reflect the categories available, who supplied the information, how the question was asked, or the context in which it was collected.
This problem is also not unique to self-reported data. Where ethnicity is recorded by somebody else, assumptions, misunderstandings or differences in recording practices can also result in a person's ethnicity being recorded differently.
For researchers using or linking administrative data, these differences have practical implications. If ethnicity is recorded differently across, for example, a person's health and education records, researchers need to decide how those different values should be interpreted or combined. There may not be one simple, universally “correct” value.
Understanding ethnicity data
As I have hoped to introduce in this blog, there is a long journey between how somebody understands their own identity and the ethnicity variable a researcher eventually sees in a dataset. Behind every variable is a process through which complex human identities have been translated into data through a series of choices, contexts and systems.
For administrative data researchers, understanding how that journey happened – who reported the information, what options they were given, when and why it was collected, and what may have been gained or lost in the process – is important when interpreting findings using ethnicity data.
In my next blog, I'll explore these questions through the example of the Windrush generation. I'll look at how broad labels can bring people together while also concealing important differences, and what the experiences of the Windrush generation can teach us about the relationship between identity, administrative records and people's lives.