A congressional investigation estimates broker breaches have cost consumers $20 billion in identity theft. Major brokers now promise to make it easier to opt out of their databases.
Sen. Maggie Hassan kicked off an investigation into data brokers in response to CalMatters reporting. Hassan speaks during a Senate Finance Committee on Capitol Hill, March 14, 2025.
Ben Curtis, AP Photo
Breaches at data brokers have cost American consumers more than $20 billion, Congress’s Joint Economic Committee revealed Friday as part of an investigation triggered by The Markup and CalMatters.
The estimated losses stem from identity theft linked to just four recent data breaches involving major brokers, the committee said in a report.
Released by the committee’s Democratic minority, the document repeatedly cited reporting into data brokers from The Markup and CalMatters, done in collaboration with WIRED.
The committee followed up directly on the Markup and CalMatters’ reporting, which in August showed how data brokers were hiding from search engines legally-mandated pages where Californians can request that the brokers delete or stop selling their data.
Shortly after that story was published, New Hampshire Democratic Sen. Maggie Hassan, ranking member of the committee, sent a letter pressing some brokers to explain their practices. In response, the report revealed, four major data brokers engaged with congressional staff and changed their practices to make it easier for consumers to control the use of their data.
Data brokers and the ‘no-index’ tag
Data brokers, as defined in the California law that requires brokers to provide consumers the so-called “opt out” pages, are companies that gather data on consumers, then sell that data to other companies who do not have a direct relationship with the consumers. Typically, companies buy such information from data brokers for marketing purposes.
Brokers can gather the data from information like public records, or more invasive methods like tracking online activity. Though brokers hold potentially sensitive information on consumers, many Americans are unaware they exist.
Under the California law, data brokers that reach a certain size are required to register and provide a clear way for consumers to request that their information be removed, that it not be sold or that they get access to it. But The Markup and CalMatters’ August story examined how several data brokers used code called the “no-index” tag on pages where consumers could exercise their right to opt out.
The tag is used to tell search engines not to index the page, meaning the information may not be returned in search results. The story noted that this created a barrier for consumers looking to block brokers from using their data. Many of the data brokers quickly removed the tag as the story was published.
In response to that initial reporting, Hassan independently contacted five major brokers, asking for more information about their practices. Only one registered broker, called Findem, declined to engage with staff or change its practices, the report said.
“Following Ranking Member Hassan’s requests, most companies took action to make their opt out and other privacy pages more visible for individuals, including by removing ‘no index’ code, adding opt-out links in more prominent locations, and publishing blog content that explains how consumers can exercise their privacy rights,” the report reads. “Ranking Member Hassan welcomes these actions as supporting greater protection for consumers against scams and other harms.”
Billions in losses
The report goes on to estimate the potential losses incurred by consumers because of recent data broker breaches, pegging the number at $20.8 billion.
Congressional staff found that hundreds of millions of people were exposed by just four major data broker breaches in the last 10 years. The breaches counted were a 2017 Equifax incident, impacting 147 million people, as well as others involving Exactis in 2018, 230 million people, National Public in 2023, 270 million people and TransUnion in 2025, 4 million people.
Using estimates of the number of people who experience identity theft after breaches, as well as an assumed median loss of $200 from thefts, the report arrived at the nearly $21 billion figure.
The report calls for action to prevent such losses in the future, including by filling gaps exposed by The Markup and CalMatters’ reporting on brokers.
“The Committee’s findings underscore the need for clear, easy access to opt-out options and more rigorous oversight within the data broker industry. Especially given the Committee’s calculation that U.S. residents have lost more than $20 billion in recent data breaches, additional action is needed to protect Americans from scams connected to data brokers,” the report reads. “At a minimum, opt-out options should be easy to locate and use.”
A legislative battle is under way over gaps that allow companies to collect and sell students’ personal information.
Students in a classroom in Sacramento on May 11, 2022.
Miguel Gutierrez Jr., CalMatters
For every aspect of a student’s life, there’s a tech company trying to digitize it. Inside the classroom, online tools proctor exams, create flashcards and submit assignments. Outside, technology coordinates school sports, helps bus drivers find the right route and maintains students’ health records.
California has a number of laws aimed at protecting children’s data privacy, but those laws have exceptions that allow many tech companies to continue packaging and selling students’ personal information.
This year, Assemblymember Dawn Addis, a San Luis Obispo Democrat, is carrying a high-profile state bill that would add new protections for students. She says it’s important, especially as the Trump admin is trying to collect data about California residents’ immigration status, gender identity, and their use of certain public benefits.
Historically, California has been a leader in data privacy. In 2014, California passed a landmark student privacy law that prohibited technology companies from selling students’ data, targeting students in advertising, or disclosing their personal information. Then in 2018, the state passed another unprecedented bill that required all companies give California users certain privacy rights, such as a chance to opt out of data collection and delete some of their information.
But as technology evolved and proliferated, privacy laws repeatedly fell short in protecting California’s students — at the same time that the federal government has tried to collect increasing amounts of personal information, Addis said.
Her bill would restrict how AI companies use student data and create new data protections for college students. Some of Sacramento’s most powerful players are paying close attention to the measure, including the California Labor Federation, which supports the bill, and the California Chamber of Commerce, which opposes it. Combined, these two groups spent nearly $8 million on campaign donations to state legislators or other political activities in 2024, according to the CalMatters Digital Democracy database. TechNet, a trade association that represents many of the most powerful tech companies, also opposes the bill.
The proposal, Assembly Bill 1159, would close certain loopholes in the state’s 2014 education privacy law, but experts say it may not be enough to prevent companies from selling students’ data.
A privacy expert struggles to keep her information private
Jen King is a privacy and data policy fellow at Stanford’s institute for AI, where she studies the tricks that companies use to gather users’ data and prevent them from opting out, sometimes known as “dark patterns.” In her personal life, she’s vigilant about avoiding online data tracking and maintains a landline in her Bay Area home to avoid giving out her cell phone number.
King doesn’t want her children’s information available online or for any company to sell, though sometimes it happens before she can stop it.
In the fall, King got an email about a platform called TeamSnap, which her 12-year-old son’s cross country coaches were using to manage the team’s roster. The company wanted her information, including her name, date of birth, gender, email address, and phone number. Once she logged in to the platform, she could see some of her son’s information, such as his name, email, and date of birth, were already listed. Photos and personal information from all of her son’s teammates were also available for her to see.
“I was super irritated,” she said. “You don’t need my birth date — I’m a freaking parent.” She acknowledged some personal information could be useful for a coach but said that other questions seem designed to help the platform sell information to data brokers and ultimately, to advertisers.
Her 17-year-old son’s data is also on TeamSnap, she later learned, because his robotics team uses it. This month, when King tried to show The Markup and CalMatters her TeamSnap account, a pop-up appeared, asking her if the company could track her activity across other apps and websites.
Federal law requires companies to get parental consent before knowingly collecting or selling data from children 12 and under, but once a child turns 13, their data is generally treated much like an adult’s information, especially when that child is interacting with tech platforms outside of school. TeamSnap’s privacy policy says it doesn’t knowingly collect personal information about users under 13 “without express parental consent,” though it says in some cases a team or organization may provide information on behalf of the child.
The policy also says that TeamSnap has “not sold the personal information of any consumer for monetary consideration” in the last 12 months, but that its “use of cookies and other tracking technologies may be considered a sale of personal information under the CCPA (California privacy law).” Information sold to advertisers and marketers included users’ names, contact information, purchase history and geolocation, the policy says.
California privacy law specifically requires certain large for-profit companies to get consent to collect data from anyone under 16. Often, consent happens when a user first opens a website and a pop-up appears, asking if the website can sell your data or track your cookies.
If a teacher, coach, or other authority figure tells a student that they have to use a website or an app, then the student cannot realistically opt out, King said. They may be too young to understand how to opt out, she added. “Most 15-, 16-year-olds don’t have any idea what this is about.”
Even older college students may have little agency in the technology they use, especially if it’s required for class or residential life. At Stanford, for example, King said her undergraduate students are often required to create Facebook accounts for student groups.
The same is true for parents. King said she reluctantly gave TeamSnap her personal information, including her name, email, date of birth, and the landline number for her home, because it was the only way to get updates about her son’s team.
How companies get around California’s education privacy laws
In 2014, California became the first state in the country to regulate education technology companies directly, but being first comes with its drawbacks. “We didn’t have examples of what best practice was,” said Amelia Vance, the president of the Public Interest Privacy Center, a nonprofit organization. The law only applies to products that “primarily” serve K-12 schools and that are designed and marketed for students.
Many tech companies argue that their products aren’t primarily intended for students or at least that they were not designed or marketed that way. The language-learning app DuoLingo, for example, has a version for schools, but the app is also popular for adults. Apps or technologies serving extracurricular programs or sports teams can claim they weren’t designed and marketed for the classroom, or that their use isn’t mandatory, said Vance. “You have this sort of black hole where there haven’t been protections.”
Addis’ bill expands the number of education technology companies that fall under the state’s student privacy laws, but the language is murky when it comes to apps or online services used outside of class.
In the case of TeamSnap, Addis’ communications director Alexis Garcia-Arrazola said the company would “most likely” fall under the scope of the bill if its technology is marketed to schools, if schools direct students to use it, and if the sports team is sponsored by the school.
Public records show that Piedmont Unified School District in Alameda County, Tamalpais Union High School District in Marin County, and Santa Monica Malibu Unified School District all purchased versions of TeamSnap, but only the Santa Monica Malibu district responded to questions about any privacy restriction imposed on the company. Brandyi Phillips, the chief communications officer for the Santa Monica Malibu schools, said the district has an annual subscription with TeamSnap, which is only available to sports staff and parents. She said there’s an agreement with the company “to protect District information and to prevent unauthorized access” but did not clarify if that agreement prevents the district from selling students’ information.
Berkeley Unified School District, where King’s children attend school, did not respond to questions about any contracts, purchase orders or agreements with TeamSnap.
Locally, school districts and colleges have the power to negotiate the privacy terms of any contract they make with a technology company, but many websites and apps offer free versions that a teacher or coach might recommend without getting formal approval from their district.
Last year, the California State University system signed a nearly $17 million contract with Open AI, the company that operates ChatGPT, including an agreement that the company will not train its models on student data. Advocates for Addis’ bill say the same privacy restrictions should apply to any AI company with access to California student data, regardless of whether the company has an agreement with the student’s school district or college.
Are privacy laws getting stricter or looser?
Addis’ bill comes as privacy laws in California and across the country are in flux. In 2020, California voters approved a proposition to create a new state agency to enforce data privacy rules and regulate the businesses that collect data. Advocates for the proposition contributed over $6.7 million to the campaign, compared to just over $50,000 contributed by the opposition, according to state data. The state agency that the proposition formed, now known as CalPrivacy, released new rules this year, restricting the use of automated decision-making technology, such as the use of AI to make admissions or hiring decisions. Those rules were originally stricter but businesses, lawmakers and Gov. Gavin Newsom pressured the CalPrivacy board to water them down.
In Washington D.C., Congress is considering changing federal law to limit how companies interact with children under 17. Separately, Congress is considering a bill that would require social media companies to prevent and mitigate children’s sexual exploitation, bullying, and self-harm. California Attorney General Rob Bonta is concerned that one version of the social media bill contains language that could erode existing protections in California law.
Bonta’s office is responsible for enforcing many of the state’s existing privacy laws. In November, he said the state worked with Connecticut and New York to reach $5.1 million in settlements against Illuminate, an education technology company that uses data to track and evaluate students’ progress, such as their testing scores and developmental milestones. The company had a data breach, exposing “sensitive information” from over 434,000 California students, the state attorney general’s office said in a statement.
It was the first time California successfully went after a company for violating the state’s landmark 2014 education privacy law.
To increase enforcement, Addis’ bill contains a new provision — the right for students and parents to sue tech companies in certain cases for privacy violations. Business and technology groups have opposed the bill, arguing that the new regulations and the right to sue would stifle investment in AI-powered learning tools.
King said that giving consumers the right to sue is often the only way to increase enforcement. Otherwise, the onus is on individual consumers to find concerning practices and try to opt out.
Despite being an expert in data privacy, King said that she struggled at first to figure out how to delete her TeamSnap account, only later to discover that she needed to send an email to the company. She laughed at the irony, since it’s these kinds of dark patterns in user design that fuel part of her research.
In academia, the strategy of trapping customers is sometimes called the “roach motel,” she explained, a reference to a popular television ad from the late 1970s for a cockroach trap.
“You can check in,” she said, “but you can never check out.”
We’ve updated Blacklight, our popular privacy tool, to check for TikTok and X trackers.
Gabriel Hongsdusit
Since 2020, readers have used Blacklight, our pioneering website privacy inspector tool, to run more than 18 million scans. Previously, Blacklight detected tracking pixels from Google and Meta. Today, we’re announcing that it can scan for two more digital trackers: TikTok and X pixels.
Scan a website
A tracking pixel is a small piece of code added to a website that sends information about the site’s users to the platform that operates the pixel. That can include details of a user’s activities, such as their browsing activity, purchases and searches. A website that embeds a pixel often does so to inform its advertising campaigns on the platform that create the pixel. When its pixel is embedded across many websites, the platform can compile a user’s data to build a detailed profile of their interests, behavior and other personal information. These profiles allow other businesses to buy ads from the platform to target categories of users — though this data can also be used for other purposes.
When you look up a website in Blacklight, it will now report if it finds the TikTok pixel or X pixel. More detailed information about the specific data being passed through pixels is also available by clicking on “Learn more” in the top right of the results, then clicking the link to “download an archive.”
To develop these new features, we partnered with a group of computer science students in Brandeis University’s Capstone in Software Engineering course. These students – Yiyou “Felix” Fan, Jiawen “Zena” Hu, Hengye Li, Hongchen “Steven” Yang and Yiquan “Frank” Zhang – researched and developed the features with the support of our product team.
Blacklight’s pixel detection features have already powered our Pixel Hunt investigations, which revealed that sensitive personal user information was being shared from government websites with Meta and Google, leading to lawsuits, removal of pixels from sites and increased government scrutiny. These new features give a fuller picture of the digital privacy landscape by exposing tracking pixels from two more companies.
We hope these new features will help you better understand what happens to your data as you navigate the internet. While Blacklight can’t say exactly what companies like TikTok and X do with our data, it can provide a starting point for deeper investigation into how that data is stored, shared and used across the web.
Do you have questions, suggestions or need help understanding your Blacklight results? You can always reach us at blacklight@themarkup.org.
New York mayor says terminating the ‘unusable’ bot will help close a budget gap
New York City Mayor Zohran Mamdani speaks at a press conference at Gracie Mansion in New York City, on Jan. 12, 2026.
Michael M. Santiago, Getty Images
This article is co-reported with THE CITY, a non-profit newsroom that serves the people of New York. Sign up for its newsletter, The Scoop.
In a press conference this week on New York City’s $12 billion budget gap, Mayor Zohran Mamdani zeroed in on the previous administration’s artificial intelligence chatbot as one of “a number of different things we’re going to pursue for savings.”
The chatbot, which was released by the Eric Adams administration in fall of 2023, was meant to provide business owners with an accessible way to check city rules and regulations. But as first documented by The Markup and THE CITY, the bot provided answers that, if followed, would lead to illegal behavior by businesses, like taking a cut of employees’ tips.
A spokesperson for the mayor, Dora Pekec, confirmed in a text message that the new administration plans to take down the chatbot. She said a member of the Mamdani transition team had seen reporting on the bot from The Markup and THE CITY and presented it to the mayor as a possible place to save funds.
At the press conference, Mamdani blamed Adams for the budget shortfall, saying he had been handed “a poisoned chalice.” To close the deficit, he said he would raise taxes on the wealthy and corporations and look “under the hood” of the city’s budget for potential savings.
When pressed by reporters on what he might cut, he singled out the chatbot.
“The previous administration had an AI chatbot that was functionally unusable,” Mamdani said. “It was costing the administration around half a million dollars. That, in and of itself, is not something that can bridge this kind of a gap, but it’s an indication of the ways in which money has been spent while refusing to account for the actual costs of what these programs are.”
The bot, built using Microsoft’s cloud computing platform, was part of an ambitious overhaul of digital services in New York called MyCity. The project was meant to streamline access to government but was criticized for relying on outside contractors.
It wasn’t clear how much it cost to maintain the chatbot. Just building the bot’s foundations reportedly cost nearly $600,000, close to the figure Mamdani provided. Pekec said they didn’t yet have a date for taking down the bot.
A broken bot
Testing by The Markup and THE CITY in 2024 showed that, despite promises from the Adams administration, the chatbot would confidently provide incorrect and potentially harmful information to visitors, even on high-stakes topics.
When asked about housing policy, for example, the bot suggested landlords could discriminate against tenants with Section 8 vouchers. Despite being an intended resource for business owners, the bot didn’t know the minimum wage, and told users it was fine to refuse to accept cash for payment despite a city law to the contrary, enacted in 2020.
After The Markup and THE CITY’s initial report was published, readers continued peppering the bot with sometimes farcical questions, which it continued failing to answer properly. The Adams administration defended the bot, saying it would improve over time.
“We’re identifying what the problems are, we’re gonna fix them, and we’re going to have the best chatbot system on the globe,” Adams said at a press conference. “People are going to come and watch what we’re doing in New York City.”
City administrators soon added disclaimers to the bot advising users to “not use its responses as legal or professional advice.” They also improved some of the bot’s answers, but also appeared to limit the kinds of questions the tool was willing to answer.
Today, the bot advises visitors to “ask an NYC government question only” and cautions that “responses may occasionally produce inaccurate or incomplete content.” Visitors must agree to accept the bot’s limitations before using it.
Update (Feb. 4, 2026): The city has taken down the chatbot, writing on the bot’s web page that its “beta test has ended” and directing visitors to NYC.gov for government information.