इस प्रोजेक्ट के बारे में
Vouch एक फाइनेंस-ग्रेड tool-calling एजेंट है जो SEC फाइलिंग्स का उपयोग करके किसी भी U.S. पब्लिक कंपनी के बारे में वित्तीय प्रश्नों का उत्तर देता है। यह dual-path architecture का पालन करता है: सटीक आंकड़े live XBRL डेटा पर deterministic tools से आते हैं (LLM वित्तीय डेटा पर कभी अंकगणित नहीं करता), जबकि narrative context 10-K, 10-Q, और 8-K दस्तावेज़ों पर pgvector dense search के माध्यम से प्राप्त किया जाता है। एक output guardrail किसी भी ऐसी संख्या को रोकता है जिसे उसके स्रोत तक ट्रेस नहीं किया जा सकता—या तो XBRL tool या उद्धृत filing passage—जिससे सिस्टम गढ़ने के बजाय abstain करता है।
मुख्य क्षमताओं में सटीक आंकड़े प्राप्त करने के tools (`get_financials`), मानक ratios (`get_ratio`), year-over-year growth (`get_growth`), custom formulas (`compute_formula`), पूर्ण statements (`get_statement`), सबसे बड़े/सबसे छोटे line items, segment/geography breakdowns, और qualitative search शामिल हैं। Ratios और growth को fixed conventions के साथ गणना किया जाता है ताकि गलत base metric चुनने की सामान्य LLM त्रुटि से बचा जा सके। Delisted या renamed फर्मों को नाम से resolve किया जाता है, dead ticker से नहीं।
सिस्टम को तीन-स्तरीय मूल्यांकन के माध्यम से मान्य किया जाता है: 84 cases का एक internal CI gate जो numeric grounding, citation, abstention, और trajectory metrics को स्कोर करता है; calibrated LLM judges (faithfulness Cohen's κ = 0.76 पर, abstaining answers पर शून्य false positives); और FinanceBench (Patronus AI) पर external benchmarking, जहाँ इसने zero-fabrication rate पर numeric set पर 94% addressable coverage और 96% narrative groundedness प्राप्त की। Adversarial red-team testing ने fabrication rate को 11% से 0 तक लाया।
उपयोगकर्ता अपने स्वयं के वित्तीय दस्तावेज़ (internal reports, non-public companies) अपलोड कर सकते हैं और एजेंट Docling के माध्यम से table-cell extraction का उपयोग करके उनसे उत्तर देता है, cell-level citations (`filename · page · row/col`) के साथ। प्रत्येक दस्तावेज़ per-user isolated और scoped है ताकि cross-file attribution रोका जा सके।
Deployment एक local service को उजागर करता है जिसमें `/ask`, `/export`, `/upload`, और एक OpenAI-compatible chat endpoint है। डिफ़ॉल्ट model o4-mini है, और सिस्टम में एक miss queue शामिल है जो भविष्य के विस्तार के लिए unhandled metrics को log करता है। Vouch LangGraph और pgvector पर बनाया गया है, जो पहले के Slug Advisor प्रोजेक्ट से retrieval और CI-eval engineering का पुनः उपयोग करता है।
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.