About this project

Civitas is an open-data AI/ML platform that aggregates data from official U.S. government sources into unified transparency scorecards for senators, House representatives, presidents, and Supreme Court justices. It also features an Action Center that surfaces trending civic issues from news analysis, auto-detects ongoing national concerns as trackable monitors, and builds a year-in-review timeline. Voting records, campaign finance, floor speeches, judicial opinions, and stated platforms are analyzed using embedding-based classification, content-based party alignment, and deterministic scoring — all running locally on a Raspberry Pi 5 with zero external API calls to cloud AI services. ## Architecture The system is built around a nightly batch pipeline that fetches data from multiple government APIs (Congress.gov, FEC, GovInfo, Senate.gov, Oyez, BLS/BEA, etc.), transforms it, analyzes it using embeddings and deterministic scoring, and persists results to SQLite databases. A FastAPI backend serves the data to a Next.js frontend. The LLM (LFM2.5-1.2B-Instruct) runs locally via llama.cpp for tasks requiring natural-language synthesis, such as Action Center issue generation and justice profile summaries. ## Key Features - **Transparency Scorecards**: Scores for senators, representatives, presidents, and justices based on voting records, campaign finance, floor speeches, and other data. - **Action Center**: Surfaces trending civic issues from news analysis, auto-detects national concerns as trackable monitors, and builds a year-in-review timeline. - **Hybrid Search**: Explore index combines embedding-based semantic search with FTS5 keyword search and citation-graph PageRank. - **Fully Local**: All models, databases, and services run on-device (Raspberry Pi 5, 16GB RAM). No cloud GPU, no third-party AI APIs. - **Deterministic Scoring**: Phase 3 analysis is fully deterministic with no LLM calls, using embedding-based classification and kNN. - **Caching Architecture**: Three independent caching systems (ApiCache, AnalysisCache, LearnedClassification) for replay, performance, and audit trails. ## Pipeline Phases The nightly pipeline processes every senator and House representative through seven phases: FETCH, TRANSFORM, ANALYZE, EXPLORE, JUSTICES, PRESIDENTS, and FINALIZE. It runs as a chain of five pipelines (Senate, Supplementary, House, Stock trades, Election) where each link starts only if the previous one finished. ## Classification Strategy Classification decisions are made using a tiered strategy: exact matches, sentence-transformer embeddings, kNN, and finally LLM for tasks requiring natural-language synthesis. The system emphasizes explainability and reproducibility, with all model weights pinned. ## Hardware Raspberry Pi 5 (16 GB RAM), NVMe SSD. All models, databases, and services run on-device.