Corpus graph — public-record entity index

POSTGRES 17.6 · PGROUTING 3.4.1 · POSTGREST

An Australian council's minutes, committee papers, contracts registers and tender notices (2015-2026), plus seven US federal references, extracted to text and indexed as a graph database in ordinary Postgres. The counts above are live from the database, and they are the corpus. Entities are people, organisations, checksum-validated ABNs, and citations. Edges are co-occurrence within a 400-character window, dated by their document, so every query has an as-at form. Traversal is a recursive CTE, shortest path and components are pgrouting, and the whole read surface is PostgREST over one edge table.

LOADING

Source corpus

what got extracted from each document
sluggenresourcetextratioentitiesmentions

The ratio is not a constant. A positioned form yields 0.048; dense regulation yields 1.089, because a PDF already compresses its own content streams. Source size does not predict database size.

Cross-document entities

0 entities appear in 2+ documents
only edges whose document existed by the date count
kindlabeldocsmentionsfound in

The discovery surface for a first-time user: entities that bridge documents are the natural entry point. Bridging is corpus-shaped: 20 of the 1521 US citation entities span documents, against sitting councillors who span 86-93 of the 103 council documents.

Document search

keyword search over extracted text

Enter a search term above.

Entity search

find a control, person or organisation by name
kindlabelmentionsdocsscore
NO MATCH

Connected components

whether the corpus separates into topics on its own
sizelargest members

The corpus separates into topical clusters on its own: security controls in one component, statutes and public laws in another, securities regulation in a third.

Selected entity

none

Select a result with OPEN.