Transparency
What we publish and why
Transparency is not a legal obligation we fulfil minimally. It is a design principle. We publish what we know, including what we do not know and what we got wrong.
Publications
What we make public
Model Cards
Every model we ship has a published model card documenting its intended use, evaluation results, known limitations, and out-of-scope uses. Model cards are updated with each significant checkpoint. They are written to be readable by non-technical stakeholders, not just researchers.
Training Data Provenance
We publish a high-level data provenance report for each major model, covering the categories of data used, the sources and their licensing status, and the filtering and curation methodology. We do not publish raw datasets, but we publish enough to allow meaningful scrutiny of our data practices.
Evaluation Results
We publish benchmark results on standard academic evaluations and our own internal evaluation suite, including low-resource and multilingual benchmarks that most labs do not report. We include disaggregated results across language, domain, and demographic group where available. We do not cherry-pick benchmarks.
Known Limitations
We maintain and publish a known limitations register for each deployed model, documenting failure modes we have identified, the conditions under which they occur, and what mitigations we have or have not yet deployed. This register is updated on a rolling basis as new limitations are discovered.
Reasoning
Decisions we've made and why
We believe that explaining the reasoning behind significant decisions is as important as the decisions themselves. Where we have made choices that have meaningful implications for users, researchers, or society, we explain why we made them and what alternatives we considered.
This includes decisions to restrict certain capabilities, to refuse certain deployment contexts, to prioritise certain languages in evaluation, and to set our Responsible Scaling Policy thresholds where we did. These explanations are published in the transparency log below and updated as decisions evolve.
Log
Transparency log
A dated record of significant decisions, policy changes, and disclosures.
April 2026
PolicySattam.ai access restricted to licensed legal practitioners
We decided to limit early access to Sattam.ai to licensed advocates and law clerks rather than making it broadly available. Our reasoning: legal information provided to unlicensed users without appropriate context creates a meaningful risk of harm in high-stakes legal situations. We will expand access as we develop appropriate safeguards for non-practitioner use.
February 2026
SafetyClark CL-2 safety evaluation results published
We published full results from our CL-2 capability and safety evaluations, including red-team findings and the two areas where Clark performed below our internal threshold on first assessment: persuasion resistance in regional language contexts and factual accuracy on contested historical events. We describe the mitigations applied and their measured effectiveness.
January 2026
GovernanceAI Constitution v1.0 ratified, public comment period summary published
After a 90-day public comment period, we ratified Version 1.0 of our AI Constitution. We received 47 substantive comments from researchers, civil society groups, and legal scholars. We publish a full summary of comments received, our response to each, and where comments led us to revise the document. Three comments led to material changes.
November 2025
DataTraining data audit: three sources removed after licensing review
An audit of Clark's pre-training corpus identified three data sources whose licensing terms we concluded did not clearly permit use in commercial AI training. We removed those sources from the corpus and conducted a targeted fine-tuning pass to reduce any residual influence. We publish the names of the removed sources and our legal reasoning.