Sources

7 sourcesCounted from our tables

44B measures almost nothing itself. It gathers, joins, and keeps — and the gathering comes from seven outside organizations, each of which does the hard part. This page says who they are, what we take from each of them, and how many records that accounts for. Every figure below is counted from our own tables rather than typed in by hand.

Artificial Analysis

artificialanalysis.ai

What it is

An independent benchmarking service that measures models itself — capability, latency, throughput, and price — and publishes the composite indices the field has largely standardized on.

What 44B ingests

Every model and creator it tracks, with the full metric set behind the model pages and the leaderboard, plus its media-model ratings. Ingested into our own tables; nothing here calls their API at request time.

586models56creators
497media ratings

Seen onModelsLabs

What it is

A metadata and benchmark aggregator: what a model is licensed for, whether its weights are open, which evaluations it has been run on, and who serves it.

What 44B ingests

The license and openness layer on every model, the benchmark catalog with its scores, the TrueSkill rankings, and the provider roster behind the /providers section.

Seen onBenchmarksProvidersModels

Semantic Scholar

www.semanticscholar.org

What it is

The academic search engine and citation graph built by the Allen Institute for AI, covering the literature and how it cites itself.

What 44B ingests

Abstracts, authors, venues, and citation counts for the Library and for the paper behind each benchmark; and a crawl across safety, evaluation, governance, alignment, interpretability and risk, whose papers are Library documents in their own right.

Seen onLibraryBenchmarks

What it is

The open-access preprint archive operated by Cornell University, where nearly all AI research appears first.

What 44B ingests

A running feed of cs.AI, cs.LG, and cs.CL, refreshed daily by cron. Metadata only — the reader embeds the paper from arXiv rather than rehosting it, and every record links back.

Seen onLibrary

Metadata used under arXiv’s terms of use; full text stays on arXiv.

Hugging Face

huggingface.co

What it is

The public repository for machine-learning models and datasets, and the home of the open safety datasets the incident and harm layers are built from.

What 44B ingests

The Butterfly Labs AI incident dataset — itself an aggregation of the AI Incident Database, NVD CVE, MITRE ATLAS, and around sixteen research, advocacy and news feeds — filtered to real incidents. The registries are listed as they arrive; feed items are listed only when they document an actual failure or harm event, because that dataset mixes announcements, papers and podcasts in with the incidents. Plus the harm-taxonomy and measurement datasets (BeaverTails, PKU-SafeRLHF, TruthfulQA, JailbreakBench) behind the behavior-measurement layer.

2,568incidents on record
7upstream feeds
4safety datasets
7,017classified rows

Seen onIncidents

Incident data CC-BY-4.0; each dataset keeps its own license on its card.

What it is

The AI Governance and Regulatory Archive, kept by Georgetown’s Emerging Technology Observatory: the laws, regulations, executive orders, standards and corporate commitments that actually govern AI, read and annotated provision by provision.

What 44B ingests

Every document in the archive, with ETO’s own summary as the abstract and the annotated provisions as the sectioned text you can read here. Each one carries its issuing authority, its status — proposed, enacted, or defunct — and the governance taxonomy ETO applies to it.

1,229documents in the Library
10,724annotated provisions
6,600governance tags applied

Seen onLibrary

Emerging Technology Observatory AGORA dataset, used under CC BY-NC 4.0 — attribution required, non-commercial use.

What it is

The U.S. Patent and Trademark Office’s own machine-learning read of its entire corpus: every U.S. patent document published since 1976, scored for whether it contains artificial intelligence and in which of eight components — machine learning, natural language, vision, speech, knowledge representation, planning, evolutionary computation, and AI hardware.

What 44B ingests

Every granted patent the models flag as containing AI, with its component scores, joined through the USPTO Patent Assignment Dataset to the organizations recorded as owning it — and from there to the labs on these pages. Current cut: the AIPD 2023 vintage (patent documents through 2023).

1,404,957AI patents
1,139,473with a recorded owner
172,153matched to a lab

Seen onPatents

U.S. Government work, no copyright restriction. Cite Pairolero et al., “The artificial intelligence patent dataset (AIPD) 2023 update,” J Technol Transf (2025), and Graham, Marco & Myers, “Patent transactions in the marketplace,” J Econ Manage Strat 27:343–371 (2018).

Each source is credited and linked wherever its data appears, not only here. Terms of art used across the site are defined in the glossary. Something we should be carrying? Tell us.