How data brokers collect and sell your personal data
Nobody signs up with a data broker. That is the defining feature of the industry: its subjects are not its customers, have usually never heard of the companies holding files on them, and receive nothing in exchange. The trade is legal, large and mostly invisible, and it runs on information you did hand over, just to somebody else.
Where the data originates
The raw material comes from four broad channels, and only one of them involves anything resembling a choice.
Public records. Property deeds, court filings, voter registration, business registrations, professional licences, marriage and bankruptcy records. This material is public by design, for reasons that predate the ability to aggregate it instantly. Individually it is unremarkable. Compiled, it produces a detailed biography.
Commercial transactions. Loyalty programmes, warranty registrations, magazine subscriptions, charitable donations, catalogue orders. The disclosure that permits onward sale is usually present in a privacy policy, phrased broadly enough to cover partners and affiliates without naming them.
Online behaviour. Advertising identifiers, cookies, app telemetry and location signals collected by software development kits embedded in ordinary apps. A weather app or a game may carry several such kits, each reporting to a different company, which is why location data circulates so widely.
Other brokers. Data sets are bought, licensed, merged and resold repeatedly. This is why correcting a record in one place rarely propagates, and why the same error follows people around for years.
How profiles get assembled
Individual records are close to worthless. Value comes from linkage, and linkage is the industry’s actual technical product.
The process works by matching identifiers across sources. An email address appears in a loyalty database and an app registration; a home address appears in a property record and a catalogue order; a hashed phone number ties an advertising profile to an offline purchase history. Each match increases confidence, and the merged record becomes more sellable than the sum of its parts.
Two products emerge. The first is the identity graph, a map of which identifiers belong to the same person across devices, addresses and accounts. The second is the segment, an inferred category such as expecting a child, recently divorced, likely to be managing a chronic condition, financially stressed. Segments are the commercial output, and they are inferences rather than facts, which matters because inferences can be wrong while still being acted upon.
That distinction is the source of most of the harm in this system. A person who is not actually pregnant, not actually in financial difficulty and not actually ill can still be priced, targeted and screened as though they were, with no visible mechanism to contest a classification they cannot see.
The full path, from origin to saleable product:
| Layer | What enters | Did you agree to this? |
|---|---|---|
| Public records | Property deeds, court filings, voter rolls, licences, bankruptcies | No agreement exists; public by design, long before aggregation was possible |
| Commercial transactions | Loyalty schemes, warranty cards, subscriptions, donations, catalogue orders | Technically yes, via a privacy policy naming “partners” rather than companies |
| Online behaviour | Ad identifiers, cookies, app telemetry, location from embedded SDKs | You permitted the app, not the third party collecting inside it |
| Broker to broker | Existing sets bought, merged and resold repeatedly | No, and no notification at any point |
| Product one: the identity graph | Which identifiers belong to the same person across devices and addresses | The technical achievement that creates the value |
| Product two: the segment | Inferred categories such as financially stressed, expecting, managing a condition | An inference sold as a fact, invisible and uncontestable |
The two bottom rows are what is actually bought and sold. Everything above them is raw material, and the shift from fact to inference happens at the last step, which is also the only step with no legal notice attached to it.
Who buys it and why
The buyers are more ordinary than the practice suggests.
Advertisers and their intermediaries buy segments to target campaigns and to measure whether an online advertisement produced an offline purchase. Retailers and insurers buy data to model risk and price. Landlords, employers and lenders buy screening reports, which in the United States pulls them into a regulated category with rules attached, a distinction brokers work hard to stay outside of. Debt collectors and process servers buy location and contact information. Political campaigns buy segments for canvassing.
Government agencies are also customers, and this is where the arrangement draws the sharpest criticism, because purchasing information on the open market can sidestep the process that would otherwise be required to obtain it.
What opt-out actually achieves
The honest answer is: something, but less than you would want, and not permanently.
Where you have a statutory right, it has teeth. Under the GDPR you can request a copy of what is held and demand erasure. Under the CCPA and comparable state laws you can request deletion and instruct a company not to sell your information. California’s Delete Act goes further by working towards a single mechanism for a deletion request across registered brokers, which is a meaningful structural improvement over sending individual letters.
The limits are real, though.
Deletion applies to the company you asked, not to the copies already sold. Public records are typically exempt, which means the spine of most profiles regenerates. Many brokers accept requests only through a form that itself asks for identifying information, which is uncomfortable and occasionally counterproductive. And the industry is fragmented enough that a thorough clearance means dozens of separate requests, repeated periodically, because new records will be collected the moment you buy a house or register a vehicle.
Paid removal services exist. They automate the sending of requests, which saves time, but they cannot exceed the rights you already have, and they add another company to the list of those holding a verified copy of your identity.
Practical removal steps
If you want to reduce your exposure rather than eliminate it, sequence matters.
Start with the people-search sites, because they are the ones that surface your home address to anyone who types your name and cause the most direct harm. Each maintains an opt-out page, generally deliberately hard to find.
Next, check whether your state or country maintains a broker registry, as California does. A registry gives you a list to work from instead of guessing, and it is the single most useful artefact in this process.
Then reduce the inflow. Turn off the advertising identifier on your phone, audit which apps hold location permission and revoke the ones that have no plausible need, and stop giving a real phone number or email to loyalty programmes that do not require one. This does not clean up history, but it slows accumulation, which over years matters more.
Finally, keep expectations calibrated. This is maintenance, not a task you finish. The relevant goal is reducing the number of companies holding a current, accurate, linkable record, and that is achievable. Disappearing is not.
Sources
- US Federal Trade Commission — enforcement and reporting on the data broker industry
- Consumer Financial Protection Bureau — guidance on consumer reporting and data rights
- California Attorney General, data broker registry — the actual list of registered data brokers operating in California, which is the closest thing to a public census of the industry
