California’s AB 2013 takes effect, forcing generative AI developers to publish training-data summaries
A new California law now requires many developers of generative AI systems made available to Californians to post a public, “high-level” summary of the training data used. The mandate, effective January 1, 2026, adds fresh compliance pressure for AI companies and could reshape how firms discuss datasets, licensing and the use of personal or synthetic data.

Generative AI companies operating in California are entering a new regulatory phase as Assembly Bill 2013 (AB 2013) is now in effect. The law requires developers of many public-use generative AI systems to publish a “high-level” summary describing key aspects of the training data used to build or substantially modify those systems. For the industry, the change raises practical questions about how to provide meaningful transparency without exposing sensitive details or inviting new legal disputes over intellectual property.

Under AB 2013, covered developers must post the summary publicly on their website, and the disclosure must address elements such as data sources, ownership, characteristics, volume, collection and processing methods, and the intellectual property status of training materials. The law also expects developers to state whether the training data includes personal information as defined by California’s privacy framework and whether synthetic (AI-generated) data was used as part of training or development.
A key detail is that the statute does not precisely define what qualifies as “high-level,” leaving room for interpretation and potentially uneven disclosures across the market. Some developers may choose broad descriptions that minimize legal risk, while others may provide more detailed inventories to demonstrate compliance and build trust. That variability could become a competitive factor, as customers and partners compare transparency practices when selecting models or vendors.
The law’s scope is also designed to be broad in who it covers: any entity that designs, codes, produces, or substantially modifies a generative AI system intended for public use in California. It applies to systems first released or updated on or after January 1, 2022, meaning a wide range of modern AI products could fall within the disclosure requirement if they are made available to Californians.
For companies, the compliance task is not just drafting text; it can involve internal mapping of dataset provenance, licensing terms and data handling procedures across multiple vendors and training runs. Large model training pipelines can combine web-scale corpora, licensed content, user-provided data, and curated datasets, and those ingredients often change with each new model version—making ongoing documentation an operational obligation rather than a one-time filing.
AB 2013 is likely to intensify a broader conversation about AI accountability and dataset governance. Even though the required disclosure is framed as a summary, it creates a public artifact that journalists, researchers, regulators and litigants can compare against other claims companies make about safety, bias, privacy and copyright. In practice, this may pressure developers to improve recordkeeping and standardize how they describe training sources, both to reduce legal exposure and to reassure customers in a market increasingly sensitive to transparency.