# Why retrieval-only chat tools fall short
When you manage hundreds or thousands of vendor contracts, the usual approach of uploading PDFs to a chat-based search tool can give fast answers for single documents but fails for portfolio questions. Those tools typically break text into indexed chunks and retrieve only the top-k matches for a query. That works when the answer lives in one chunk, but it misses aggregation tasks like total contract value, counts of expiring agreements, or the most expensive contract across the whole portfolio.
# The practical insight: structured extraction, then analytics
# Pipeline overview
- Store contract PDFs in Amazon S3. Each upload triggers the processing pipeline.
- An extraction agent reads the PDF and extracts eight key fields, each with a confidence score.
- A verification agent independently extracts the same fields to cross-check results.
- If the two agents disagree on signature detection, a computer-vision service (Amazon Textract) performs a deterministic tiebreaker.
- Verified results are stored in Amazon Aurora PostgreSQL for analytics.
- The original PDFs remain in a knowledge base for single-contract lookup and detailed reading.
# Application and user experience
A single React web application provides the interface. Users can:
- Run portfolio queries via embedded dashboards powered by Amazon Quick analytics to get instant aggregates (total exposure, expiring contracts, highest-value agreements).
- Ask natural-language questions for single-contract details that reference the stored PDFs and the verification trail.
- Review confidence scores and verification outcomes when a field matters for compliance or negotiation.
# Verification and accuracy strategy
Relying on a single automated extractor can produce inconsistent results at scale. This design uses two independent models to extract and cross-verify fields. When the extractors disagree on signatures, the system defers to a deterministic computer-vision result. This layered approach reduces false positives and gives teams a clear audit trail showing which fields were machine-extracted, double-checked, and resolved by vision.
# Why this architecture scales
- Databases and analytics engines are designed for aggregation. Once key fields are structured, adding new contracts becomes an incremental ingestion task rather than repeated manual work.
- The separation of concerns keeps the knowledge base for detailed retrieval and the relational database for math, eliminating retrieval-tool blind spots on portfolio questions.
# Where to focus next