Research tools / Reproduce
Data downloads
Take the underlying files into your own research workflow. Each export includes its schema; the citation index connects recorded facts back to their sources.
The snapshot includes 125,349 source citations and 1,938 programs in the FY2026 index. Datasets marked “cited” include row-level citation linkage; others are citation-tier pending (see methodology).
Schema & data dictionary · source linkage
Schemas & data dictionary: every Parquet file embeds its column schema (readable via DuckDB DESCRIBE), and the data explorer lists each dataset with row counts and a description of what one row represents. Keep the citation index alongside your data export so fact IDs remain traceable after a join or calculation.
Bundle built: . Files are in Apache Parquet format, readable with DuckDB, pandas, R arrow/duckdb, or any Parquet-compatible tool. Each file includes the same provenance metadata that backs on-screen figures.
One row per (program element × amount type) figure in the FY2026 President's Budget R-1/P-1 workbooks, carrying the source sheet and cell address it was read from. Fenced to the PB2026 edition. Add rows only: the P-1 flags each row Add or Non-Add, and Non-Add rows are memo lines the exhibit does not add into its own totals, so summing the workbook without that filter double-counts. P-1R rows are the reserve-component subset of P-1 lines and are never summed with them.
budget_lines.parquetThe decade sibling of budget_lines: one row per (President's Budget edition × program element × amount type) figure across all ten editions PB2017–PB2026, each cited to its own edition's workbook cell. Editions are parallel publications, never reconciled.
One row per (program element × project × budget scenario) cost figure extracted from J-book R-2/P-40 XML, with its XML element path and source-PDF SHA-256. `account` is the appropriation the figure was filed under (NULL for R-2/RDT&E rows, which carry none): ten PB2026 budget-line codes are shared by two different Navy programs in two different appropriations, so summing this table by pe_bli alone adds those pairs together.
Page resolution is partial: of 17,900 rows carrying a non-zero amount, 3,407 resolve to a unique PDF page and 6,472 to the first page the amount appears on; the remaining 8,021 (45%) resolve to no page and cite their XML path instead.
jbook_details.parquetOne row per J-book narrative text block (mission, description, justification or accomplishment/planned-program) with its XML element path and source-PDF SHA-256. Fenced to the PB2026 edition, plus the PB2017–PB2025 narratives a program-lineage edge cites: citation targets only, never a program page's own prose.
jbook_narratives.parquetOne row per (program element × organization) with FY2024 actuals, FY2025 total, FY2026 total and the FY25→FY26 delta pivoted side by side.
fct_budget_trajectory.parquetOne row per program element that has full R-2/P-40 J-book detail (the detail-grade tier, NOT the full page universe; see the corpus statement on /data/), with org, exhibit family, project count and reconciliation status.
dim_programs.parquetOne row per contractor entity family, name-normalized across USAspending/SAM.gov UEI registrations, with its UEI count, total DoD obligation and worst match-confidence tier. The sam_* columns carry the SAM.gov Entity Management record of the family's dominant registration (status, CAGE, legal business name, business types, primary NAICS, expiry) where the bounded extract has reached it — enrichment, never an input to the confidence tier, and NULL wherever it has not.
dim_entities.parquetOne row per (contractor family × filing year) of Senate LDA lobbying totals, beside the family's DoD obligations (repeated on each year row). Income and expense are non-additive: a self-filer's expense can include its outside firms' income, so lobbying_total_usd, a plain sum, can double-count.
One row per (LDA filing × matched program element) mention — a filing appears once for every program its issue text has keyword co-occurrence with (an exact PE/BLI code, a curated alias, or >=2 distinct title words; see evidence_kind), never a claim that the filing names the program. Rows count evidence-tiered mentions, not filings.
One row per (program element × USAspending award) crosswalk link, each carrying its match method and confidence tier. A link is an inference, not a reported fact.
fct_budget_to_awards.parquetOne row per (pop_state × pop_district) place of performance, with transaction count and total obligation over every DoD award transaction the site loads that records a pop_state, not only crosswalk-linked awards. pop_state is not normalized (contracts: two-letter code; assistance: state name).
One row per (state × congressional district), with obligation dollars counted once per DISTINCT high-confidence-crosswalked award — the district headline. Its per-(district, program element, appropriation account) sibling, fct_district_programs, is not summable: an award matched to N program elements appears N times with the same dollars there.
One row per program element carrying at least one published crosswalk link, on TWO bases: *_all over every published link (high and medium confidence), *_high over high-confidence links alone. hhi_high and top_family_high are NULL below the floor: 3 linked awards, 2 families holding positive dollars (positive_family_count_high), positive program_dollars_high (ROADMAP #80).
7 of its 536 rows are code-level (scope = 'code'): a budget-line code two or more programs share. On 3 of them (0145, 3010 and 3215) more than one member's key carries links, so the row pools every member key with links and describes no single program. On 4 (2101, 2292, 3050 and 4217) one member's key carries every link, so the row is that member's figure: 2101 is 2101-WPN's, 2292 is 2292-WPN's, 3050 is 3050-OPN's, 4217 is 4217-OPN's. member_keys_with_links counts the member programs whose own key carries a published link.
One row per (jurisdiction × comparable spending category × fiscal year) in the CA/CT state-checkbook pilot, with population and per-capita amount.
fct_state_per_capita.parquetOne row per federal agency reporting improper payments, with the derived dollar exposure and weighted error rate from paymentaccuracy.gov.
fct_improper_exposure.parquetOne row per individual lobbyist named in Senate LDA filings, with covered-position and revolving-door flags and the UUID/URL of the filing that disclosed them.
dim_lobbyists.parquetOne row per source citation, keyed by fact_id, in 10 kinds: workbook (a President's Budget workbook cell), derived (a recorded value with its formula), lda_filing (a Senate LDA filing), jbook_narrative (a J-book narrative passage), jbook_pdf (a figure printed in a J-book PDF, or the PDF alone where that figure is unresolved), usaspending (a USAspending API query), announcement (a DoD contract announcement), subaward (a subaward on a USAspending prime award), state_file (a state spending source file) and state_soql (a state open-data query).
Additional assets (not in table above)
pdfs/— SHA-named J-book PDFsworkbooks/— 30 R-1/P-1 Excel rollup files