btay.io/wiki

RAG for transit planning

IFP's proposal to make past transit project records searchable, citable, and reusable.

Updated 1 h ago

Source: Institute for Progress, Use AI to Improve Transit Planning, June 17, 2026.

AI-assisted summary, condensed from the original source notes. The source's numerical claims have not been independently reproduced.

The proposal

The Federal Transit Administration would collect past project reports in one repository and give agency staff a retrieval-augmented generation (RAG) interface over it. Answers would cite the records behind them.

The intended benefit is institutional memory: planners could find comparable projects, surface recurring risks, and reuse earlier work instead of commissioning it again. The paper argues this could reduce consultant reliance and construction delays. Those are proposed outcomes, not demonstrated savings.

What the system needs

LayerPurpose
Project recordsPreserve the evidence
Shared metadataMake projects comparable
RetrievalFind relevant passages
Cited answersLet staff inspect support
Agency ownerMaintain data and quality

The initial records include oversight reports, risk analyses, environmental studies, and grant agreements. Local utility and zoning data could follow, subject to access rights and data quality. A single technical owner and a planner working group would guide the platform (source PDF, pp. 6–8).

What the evidence establishes

The paper cites costly data gaps and a Caltrans/UCLA prototype that reportedly used roughly 12,000 pages of transit records (pp. 3–5). These examples support investigating document reuse. They do not establish how much a national platform would save.

The proposal supplies no measured retrieval benchmark, controlled user study, or full cost-benefit model. Its project cost examples come from other research whose methods are not reproduced in the paper.

The original PDF also has inconsistent author attribution: its opening page names Santi Ruiz, while later artwork and biographies name Elizabeth Speed and Bennett Capozzi. This brief attributes the proposal to IFP.

What a pilot should answer

Start with one corpus and one task, such as finding earlier utility-relocation risks. Compare staff using ordinary document search with staff using cited AI answers. Measure correct answers, supported citations, time spent checking, and total task time.

Include outdated reports, conflicting records, and questions the corpus cannot answer. A confident answer with an irrelevant citation must count as a failure. AI evaluation resources lists tools for running that comparison.

This is the same maintenance problem behind an AI-assisted wiki: preserve sources, make useful knowledge easier to find, and keep a person responsible for what gets treated as true.

Related pages

On this page