/PROJECTS

Live utilities, legal corpora, datasets, and open benchmarks

These projects are designed around a simple thesis: models are increasingly commoditised, but clean, verified, provenance-rich data remains scarce. The focus is not generic dashboards. It is analyst-grade infrastructure that can be used by humans, LLMs, legal researchers, and downstream data pipelines.

LIVE DATA UTILITIES
LIVE DATA UTILITY ID: OAR · DAILY CATALOGUE INGESTION

Orbital Asset Registry (OAR)

Orbital Asset Registry is a live web registry that assigns stable identities to catalogued objects in Earth orbit, resolves active satellites to their operating companies and corporate status, and uses operators' own regulatory filings to track how and when each object is disposed of or re-enters.

  • Permanent Orbital Asset IDs for approximately 70,000 catalogued orbital objects, with roughly 13,000 payloads resolved to 34 operating companies.
  • Operator intelligence across active satellites, including sourced, versioned corporate status and distress or consolidation signals, such as EchoStar's Chapter 11 filing and the OneWeb/Eutelsat and Intelsat/SES mergers.
  • Lifecycle and re-entry intelligence: parses operators' own FCC disposal reports to reconcile 2,245 reported deorbit, re-entry, and disposal-failure events against the catalogue. It surfaces operator-admitted disposal failures and cross-checks reported re-entries against independently observed decay.
  • State-succession attribution mapping roughly 25,000 legacy Soviet-era objects to the Russian Federation as the responsible successor state, with the legal basis cited.
  • Provenance-first architecture: source, date, content hash, versioned status history, daily catalogue ingestion, and monthly operator-status or weekly re-entry-report monitoring.
REGISTRY SPEC
Catalogued objects~70,000
Payloads → operators13,000 / 34
Disposal events reconciled2,245
Legacy objects attributed~25,000
ProvenanceHASH + VERSIONED
MonitoringDAILY / WEEKLY / MONTHLY
LIVE DATA UTILITY SPACE FINANCE INSTITUTE · UPDATED WEEKLY

The Observatory

A living, automatically-updated citation map of the finance-under-physical-constraint literature, published by the Space Finance Institute. It tracks the works applying relativity, information-delay, and light-speed constraints to finance, and makes explicit the citations that should exist but don't.

  • Works mapped across five strands: economics & finance, physics, mathematics, law & policy, and practice & grey literature, spanning 1951 to the present.
  • Curated "should-cite" edges: every missing-citation link between two works carries a written justification, not just a graph inference.
  • Automated discovery, human-gated publication: daily OpenAlex sweeps, forward-citation walks, and reference-list cross-checks feed a private curation queue; nothing publishes without a human-approved justification.
  • Machine-readable exports: a live map.json and dated snapshot archive, licensed CC BY 4.0.
SPEC
Literature strands5
Coverage1951–PRESENT
Should-cite edgesCURATED + JUSTIFIED
DiscoveryDAILY OPENALEX SWEEPS
ExportsCC BY 4.0
OPEN LEGAL CORPORA
OPEN LEGAL CORPUS HF DATASET · GITHUB · SHA-256

Space Law Corpus

A neutral, provenance-first, machine-readable record of international and national space law. Every authoritative text is stored with its official source, retrieval date, citation, language, SHA-256 content hash, and dated version history. It is now also published as a HF dataset for downstream AI and legal-informatics workflows.

  • 15 verified instruments: UN space treaties, UN General Assembly principles, COPUOS guidelines, and national statutes from the US, Luxembourg, and France.
  • Authoritative legal texts are kept strictly separate from generated or derived material such as translations, concept tags, and structure.
  • Published across a public corpus site, GitHub source repository, and HF dataset.
  • Designed as neutral legal infrastructure: open, self-validating, and maintained with continuous integrity checks.
SPEC
Verified instruments15
LayersAUTHORITATIVE / DERIVED
PublicationSITE + GITHUB + HF
IntegritySELF-VALIDATING CI
OPEN LEGAL CORPUS HF DATASET · GITHUB · SHA-256

Deep Seabed Mining Law Corpus

An open, neutral, provenance-first corpus of the law governing mining the deep seabed. It covers the fast-evolving international and national legal framework for mineral extraction from the ocean floor beyond national jurisdiction, without taking a position for or against seabed mining.

  • 10 foundational instruments spanning UNCLOS, the ISA exploration regulations, the draft ISA Exploitation Regulations / Mining Code, the 2011 ITLOS Advisory Opinion, and the parallel US regime.
  • Provenance-first: official source, retrieval date, citation, language, and SHA-256 content hash for every authentic text; key documents anchored to byte-exact official PDFs.
  • Two-layer design: authoritative source texts are kept separate from generated analysis, structure, and approximately 2,300 neutral concept tags across 20 categories.
  • Self-maintaining: automated monthly monitoring of 18 official sources across ISA, UN, ITLOS, and the US Federal Register.
SPEC
Foundational instruments10
Concept tags~2,300 / 20 CATEGORIES
Sources monitored18 · MONTHLY
AnchoringBYTE-EXACT PDFS
OPEN LEGAL CORPUS HF DATASET · GITHUB · SHA-256

BBNJ / High Seas Treaty Corpus

An open, neutral, provenance-first corpus of the law governing marine biodiversity beyond national jurisdiction. It captures the 2023 BBNJ Agreement / High Seas Treaty, its UNCLOS parent framework, implementing agreements, founding UN resolutions, and the 2026 Preparatory Commission report.

  • 15 authoritative instruments, including UNCLOS, all three UNCLOS implementing agreements, the 2023 BBNJ Agreement in all six authentic UN languages, founding UN General Assembly instruments, and the 2026 Preparatory Commission report.
  • Provenance-first: official citation, retrieval date, authoritative-status flag, authentic language, and SHA-256 content hash for every text.
  • Honest fidelity: texts are flagged as extracted verified, extracted unverified, or OCR unverified rather than pretending extraction quality is uniform.
  • Machine-readable value: document and provision exports, neutral concept tags, and in-house cross-reference mapping from BBNJ provisions to the UNCLOS articles they invoke.
  • Self-maintaining: CI integrity checks and scheduled monitoring for the first Conference of the Parties in January 2027.
SPEC
Authoritative instruments15
Authentic languages6 (UN)
Fidelity flags3-TIER, HONEST
Next milestoneCOP-1 · JAN 2027
OPEN LEGAL CORPUS HF DATASET · GITHUB · ZENODO DOI

AML/CTF & Financial-Sanctions Legal Corpus

A neutral, provenance-first, machine-readable record of anti-money-laundering / counter-terrorism-financing and financial-sanctions law across ten cross-referenced jurisdictions, rooted in the UN treaty and Security Council backbone, with the FATF Recommendations as the connective spine.

  • 69 instruments, 7,127 concept-tagged provisions across Australia, the UK, the US, the EU, Canada, Singapore, Switzerland, Hong Kong, Japan, and the UAE, each recorded from its own official source, never merged.
  • FATF cross-jurisdiction crosswalk: maps each FATF Recommendation to its UN treaty / Security Council basis and to the national provisions that implement it.
  • Provenance-first: byte-exact captured originals, SHA-256 hashes, append-only dated versioning, automated CI validation, and a Zenodo DOI.
  • Honest about restrictions: copyrighted FATF material is held under a restricted posture (hash, citation, and excerpt only) rather than silently omitted.
SPEC
Instruments69
Tagged provisions7,127
Jurisdictions10
VersioningAPPEND-ONLY, DATED
ArchiveZENODO DOI
OPEN LEGAL CORPUS GITHUB · MANIFEST ROOT HASH

Digital-Asset Law Corpus

A neutral, provenance-first, machine-readable record of digital-asset market-structure and prudential law across 21 jurisdictions, captured in each instrument's authentic enacting language rather than translated.

  • 57 authoritative instruments across 21 jurisdictions, in authentic languages spanning Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Maltese, Portuguese, Spanish, Thai, Turkish, and Vietnamese.
  • Two-layer design: authoritative source texts are kept strictly separate from derived structure and concept tags.
  • Offline integrity verification: a manifest root hash over every file, checked with a dependency-free script.
  • Honest coverage: the public release ships only records with confirmed clean reuse terms; 13 further records are held back pending licence confirmation rather than published prematurely.
SPEC
Instruments57
Jurisdictions21
Authentic languages15
IntegrityMANIFEST ROOT HASH
OPEN SOURCE, DATASETS & BENCHMARKS
OPEN SOURCE APACHE 2.0 · ANTHROPIC FINANCIAL-SERVICES REPO

SEC EDGAR Securitisation Connector & Skills

A free, open-source Claude / Cowork plugin with a local SEC EDGAR data connector that turns US structured-finance filings into analyst-ready output. It searches registered ABS and CMBS deals, pulls prospectuses and investor reports, and runs loan-level analysis on Form ABS-EE data. Contributed back to Anthropic's financial-services repository.

  • Local MCP connector (Python, standard library) exposing six tools; streams 100 MB+ ABS-EE loan tapes with flat memory.
  • Loan-level analytics for auto ABS and conduit CMBS: pool stratifications, balance-weighted coupon / FICO / DSCR / debt yield, property and geographic concentration, and the maturity wall.
  • The only no-subscription connector in the marketplace: free, public, rights-clean EDGAR data. Validated on real deals; Apache 2.0.
SPEC
Tools exposed6
Loan-tape streaming100 MB+ FLAT MEMORY
Asset classesAUTO ABS · CONDUIT CMBS
LicenceAPACHE 2.0
Data costNO SUBSCRIPTION
OPEN SOURCE v0.1 SCAFFOLD · APACHE 2.0

LexProvenance MCP Connector

A free, read-only Model Context Protocol connector that lets AI agents search, fetch, and cite neutral, provenance-tracked bodies of law, with a tamper-evident SHA-256 content hash on every citation, not just a link.

  • Read-only by design: five tools (list_corpora, search_corpus, fetch_document, fetch_provision, get_citation), no write, edit, ingest, or delete tools, enforced by a dedicated test.
  • Verifiable citations: get_citation returns the exact SHA-256 hash and dated version it cites, so a citation can be re-checked later, not just trusted.
  • Corpus-agnostic: ships today with sample manifests for the Space Law, Deep Seabed Mining, and BBNJ/High Seas Treaty corpora; built to sit in front of any provenance-first corpus.
  • v0.1 public scaffold, Apache 2.0 code and schemas; the free connector layer sits in front of a separate, premium enriched-corpus offering.
SPEC
Tools5 · READ-ONLY
Citation integritySHA-256 + DATED VERSION
Sample corpora3
LicenceAPACHE 2.0
OPEN BENCHMARK BENCHMARK + ARTICLE

UK Accounts LLM Benchmark

Can frontier LLMs read UK company accounts? A 1,000-question verified benchmark and five-model evaluation comparing proprietary and open-weight systems. The core finding: LLMs can read the accounts well; the defensible asset is the clean, verified data layer.

  • Benchmark coverage across Claude, GPT-5.5, Gemini, DeepSeek V4 Pro, and GLM 5.2.
  • Designed to show where model performance ends and data infrastructure begins.
SPEC
Verified questions1,000
Models evaluated5
Contamination controlPOST-RELEASE FILINGS
DATASET PROJECT HF DATASET · PIPELINE ON GITHUB

UK Company Financials

An ML-ready dataset of 3.7 million UK company filings across approximately 3.5 million companies, parsed from Companies House iXBRL accounts and normalised across the FRS 102 and FRS 105 taxonomies.

  • Turnover, assets, liabilities, net assets, funding structure, and employees.
  • Per-figure provenance for auditability and model-training confidence.
  • Normalised UK GAAP fields, GDPR-clean design, and planned quarterly refreshes.
SPEC
Filings3.7M
Companies~3.5M
TaxonomiesFRS 102 / FRS 105
ProvenancePER-FIGURE
RefreshQUARTERLY (PLANNED)