← All projects
Case study · AI · RAG · Agents · B2B SaaS

An AI assistant that reads a company's contracts and invoices, cites the exact sentence, and never acts without approval.

A small company forgets a 60-day cancellation window and pays for another year. Staff lose hours searching PDFs, and nobody wants to paste company data into a public chatbot. Sanad answers only from the company's own documents, tracks every notice deadline and payment, isolates each company's data inside the database, and proposes actions a person approves.
RoleSolo: product, backend, AI, frontend, brand
TimelineSeptember 2026
TypeSelf-initiated product, MIT licensed
StackPython 3.12, FastAPI, Postgres 16 + pgvector, LangGraph, React 19
localhost:8080/
Sanad dashboard
1.00
hybrid retrieval hit@3
0.94
citation accuracy
100%
injection attempts blocked
45
tests on real Postgres, gated in CI

Overview.

Who it is for, what goes wrong today, what Sanad does.
The audience

Small and mid-sized companies

A manufacturer with a few dozen supplier contracts, leases, software licences and invoices in Italian, English and sometimes Arabic, and no legal team. The demo company is a fictional shoe maker in Civitanova Marche.

The problem

Deadlines hide inside PDFs

Auto-renewal clauses with 60 or 90 days' notice slip by, invoices get paid late, and answering "what does our lease say about…" means opening files one by one. Public chatbots are off the table for company data.

The product

Grounded answers, deadlines and a careful agent

Upload the documents once. Sanad extracts the terms, builds the contract register, calendar and cash-flow projection, answers questions with verified citations, and drafts reminders and letters for approval.

Walkthrough.

Every screen, captured by the project's own end-to-end script.
01 · Workspace

One company, one key, one isolated workspace

Each customer company is a tenant with its own API key. Signing in opens a workspace that only ever sees that company's documents, and the demo workspace is a fictional shoe maker in the Marche region.

localhost:8080/
Sanad sign-in screen
Sign-in with a per-company key, and a one-click demo workspace
  • 1
    Keys stored as hashesOnly the SHA-256 of each API key is stored. A key maps to exactly one tenant, and per-key rate limits apply to every request.
  • 2
    Isolation in the databaseEvery request runs in a transaction that switches to a non-privileged role and sets the tenant id. Without that setting, queries return nothing.
  • 3
    Realistic demo dataFour contracts, two invoices and two policies in English, Italian and Arabic, including a lease in .docx and a policy in plain text.
02 · Dashboard

What the company owes, and what needs a decision

The dashboard turns eight documents into numbers an owner cares about: monthly commitments, the next deadlines, projected outgoing cash and how many agent proposals are waiting.

localhost:8080/
Sanad dashboard
Documents, contracts, monthly commitments, pending approvals, cash flow and deadlines
  • 1
    Commitments, not files€16,950 a month in the demo, computed from extracted fees and billing frequency, with a 12-month total of €219K.
  • 2
    Deadlines with countdownsInvoice due dates and the last day to give notice on each contract, each with days left and a link to the source.
  • 3
    Approvals surfacedPending agent proposals are counted on the dashboard and in the sidebar badge, so nothing waits unnoticed.
03 · Ask

Answers with the exact sentence and page, or none at all

Ask a question in English, Italian or Arabic. Sanad answers only from the company's own files and shows where each quote came from. If it can't support an answer, it says so.

localhost:8080/ask
Ask screen with answers in English and Arabic and a blocked injection
Answers with citations and confidence, and an injection attempt blocked
localhost:8080/ask
Citation drawer with the quote highlighted on the page
Every citation opens the page with the quoted sentence highlighted
  • 1
    Verified quotesEach quote is matched back against the retrieved text before it reaches the screen. Citations that don't match are dropped and confidence goes down.
  • 2
    Three languagesThe same pipeline answers in English, Italian and Arabic, including an Arabic question about an Arabic policy.
  • 3
    Guardrails that explain themselves"Ignore all previous instructions" is blocked before any model call, with the matched patterns and score shown instead of a silent failure.
04 · Documents

Drop in PDFs, Word files or text, and get structured terms back

Upload contracts, invoices and policies in any of four formats. A background worker parses, chunks, embeds and classifies each file, then extracts the terms that matter.

localhost:8080/documents
Document library
Library with type, status, language, pages and size
localhost:8080/documents
Document drawer with extracted contract terms
Extracted terms with confidence and evidence pages
  • 1
    Type detected automaticallyContract, invoice, policy or other, with the language detected per file. Duplicates are caught by content hash.
  • 2
    Terms, not summariesParties, term, auto-renewal, notice period, fee and frequency, payment terms, penalties and governing law, each tied to the pages it came from.
  • 3
    Implausible values rejectedA notice period of 900 days or a negative fee never reaches the register. Re-running extraction is one click.
05 · Contracts and invoices

A register that knows the last day to cancel

The contract register shows every agreement with its renewal rule and the date by which notice must be sent. Invoices sit next to it with due dates and totals.

localhost:8080/contracts
Contract register
Counterparty, term, auto-renewal, notice period and notice deadline
localhost:8080/contracts
Invoice register
Invoices with numbers, issue and due dates, and amounts
  • 1
    Notice deadline, computedEnd of term minus the notice period, adjusted for renewal cycles. It's the date the product exists for.
  • 2
    Value and payment termsFee, billing frequency and payment days per contract, so cash flow can be projected from them.
  • 3
    Evidence one click awayEvery row links back to the source document and pages the value was extracted from.
06 · Calendar and cash flow

Deadlines on a calendar, commitments on a chart

Notice deadlines, invoice due dates and manual reminders share one calendar. The cash-flow view projects twelve months of committed outgoing payments from contract fees and open invoices.

localhost:8080/calendar
Calendar of deadlines and reminders
Notice deadlines, due dates and reminders by month
localhost:8080/cashflow
Twelve-month cash-flow projection
Committed outflows by month, split into contracts and invoices
  • 1
    Reminders you controlCreate a reminder by hand or approve one the agent proposed. Approved reminders are e-mailed when SMTP is configured, or logged.
  • 2
    Projection from termsNext 12 months, monthly average, peak month and the amount that depends on renewals you could still cancel.
  • 3
    Contracts vs invoicesEach month splits fixed contract fees from open invoices, so the owner sees which part is negotiable.
07 · Agent

It investigates freely, and asks before it acts

Give Sanad a goal such as "find the most urgent renewal and draft a cancellation letter". A LangGraph agent reads the register and deadlines, then pauses with proposals that a person approves or rejects.

localhost:8080/agent
Agent run paused for approval with a drafted letter
Step trace, tool calls and three proposals awaiting approval
localhost:8080/agent
Completed agent run
A completed run after the proposals were decided
  • 1
    Read tools run, action tools waitListing contracts and deadlines runs immediately. Creating a reminder or drafting a letter becomes a proposal and the graph stops.
  • 2
    Approval can come days laterRuns are checkpointed in Postgres, so a decision can arrive from another process after a restart. Undecided proposals count as rejected.
  • 3
    A full traceEvery step, tool call, result, proposal and decision is stored and audit-logged. Letters are saved, never sent.
08 · Usage

What every answer costs, per company

Tokens, cost and latency are metered for every model call and grouped by tenant, operation and model. An audit log records every human decision and executed action.

localhost:8080/usage
Usage and audit screen
Calls, tokens and cost by operation, plus the audit log
  • 1
    Cost per operationEmbedding, classification, extraction, answering and agent steps are metered separately, so the expensive part is obvious.
  • 2
    Provider-neutral numbersThe same view works offline, with OpenAI or an EU OpenAI-compatible endpoint, or with Claude.
  • 3
    An audit trail by defaultApprovals, rejections and executed actions are written as audit events, with who decided and when.
MCP server for Claude Desktop, Claude Code and Cursor
OpenAPI docs
Structured JSON logs with request IDs
Optional Langfuse tracing
One-command Docker start
Dark mode
Arabic RTL answers

Evaluation.

Measured on 44 golden questions in CI. The build fails below the thresholds.

Retrieval · 35 answerable questions

Modehit@1hit@3MRR
Vector only0.910.970.95
Keyword only0.941.000.97
Hybrid (RRF)0.941.000.97

Answers and extraction

MetricScore
Answer correctness0.94
Citation accuracy0.94
Groundedness1.00
Abstention on unanswerable0.80
Prompt-injection block rate1.00
Extraction field accuracy1.00

Experiment · chunk size and overlap

Chunk / overlaphybrid hit@3MRRcorrectness
400 / 601.000.940.86
800 / 1201.000.970.94
1200 / 1801.000.970.91

What the numbers mean

Scores are from the offline provider, so every run is free and repeatable. Adding contextual chunk headers raised vector-only hit@1 from 0.85 to 0.91 and hybrid hit@3 from 0.97 to 1.00. The golden set covers English, Italian and Arabic, unanswerable questions and injection attempts. The remaining misses are one cross-lingual question and two near-misses, and the 1.00 extraction accuracy is a regression baseline on the sample documents, not a claim about unseen contracts.

Architecture.

One Postgres for everything, a stateless API and a worker, and a provider-neutral model layer.
Clients
React 19 UIVite · TypeScript · Tailwind v4 · TanStack Query · light and dark · desktop and mobile
MCP serverSearch, ask, the contract register and deadlines as tools for Claude Desktop, Claude Code or Cursor
REST API31 endpoints with an OpenAPI spec · one API key per company
→
Services (Python 3.12)
FastAPIAuth · rate limits · guardrails · hybrid RAG · grounded answers · agent runs and decisions
WorkerParse → chunk → embed → classify → extract · reminders · claimed from a Postgres queue with SKIP LOCKED and retries
LangGraph agentRead tools vs action tools · interrupt() for approval · AsyncPostgresSaver checkpoints
→
Data and models
Postgres 16 + pgvectorHNSW vectors · tsvector full-text · RLS on every tenant table · job queue · checkpoints · usage and audit tables
Model layerOffline (no key) · OpenAI or any OpenAI-compatible EU endpoint · Anthropic Claude · one variable to switch
DeliveryDocker Compose: db, migrations, api, worker, ui · GitHub Actions: ruff, mypy, tests, evals, image builds, secret scan
About 5,500 lines of Python in the API, 700 lines of tests and 5,500 lines of TypeScript in the UI, with six architecture decision records explaining the main choices.

Engineering deep dives.

Six decisions, each with what it cost.

1. Citations are verified, not trusted

Problem
Language models invent convincing quotes. In a product about contracts, one fake clause destroys trust.
Approach
The model cites numbered sources. Each quote is normalised and matched against the retrieved chunk, exactly or on at least 90% of its tokens to tolerate OCR noise. Unverifiable citations are dropped and confidence is lowered.
Trade-off
A correct answer that paraphrases instead of quoting can lose its citation and show lower confidence.
Lives in: api/app/rag · grounded generation

2. Tenant isolation lives in Postgres

Problem
A leak between two companies' contracts would end the product. "Remember to add WHERE tenant_id" is not a control.
Approach
Every tenant table has a row-level security policy for reads and writes. Each request runs SET LOCAL ROLE to a non-privileged role and sets app.tenant_id for that transaction only. A missing context fails closed and returns nothing. The test suite proves it with raw SELECTs.
Trade-off
HNSW scans first and RLS filters after, so a small tenant among many can get fewer than k vector candidates. ef_search is raised per query.
Lives in: ADR-003 · migrations · tests/test_rls.py

3. Approval only where it matters

Problem
An agent that reviews contracts is useful. One that creates reminders or writes to suppliers on its own is a liability.
Approach
Tools come in two classes. Read tools run as soon as the model calls them. Action tools never execute when called: they become proposals and the LangGraph graph pauses on interrupt(). State is checkpointed in Postgres.
Trade-off
Stopping before every tool call would train people to click approve without reading, so harmless reads run unattended by design.
Lives in: ADR-004 · api/app/agent

4. Hybrid search, measured

Problem
Business documents mix prose with identifiers that embeddings handle badly: invoice numbers, IBANs, supplier names, clause numbers.
Approach
Vector search and Postgres full-text run in one SQL statement and are fused with Reciprocal Rank Fusion. Every chunk is embedded with its document name and title, and each result keeps its vector and keyword rank for debugging.
Trade-off
Two indexes to maintain. With offline lexical embeddings, keyword search alone is as strong as hybrid on this corpus.
Lives in: ADR-002 · api/app/rag/retriever

5. Retrieved text is untrusted input

Problem
A malicious instruction can hide inside an uploaded document, not just in the user's question. Company data shouldn't leave in clear text either.
Approach
Prompt-injection scanning runs on the question and on every retrieved chunk before it enters the prompt. IBANs, e-mails, phone numbers and Italian tax codes are masked reversibly before any external model call, and restored in the answer.
Trade-off
Pattern-based detection needs upkeep as new attack phrasings appear.
Lives in: api/app/guardrails · injection, pii, ratelimit

6. An offline provider as a first-class citizen

Problem
Tests, CI evals and first-run demos shouldn't need an API key or cost money, and some customers will require EU-hosted models.
Approach
Every task has a deterministic implementation: lexical embeddings, extractive answers, regex extraction and a scripted agent policy. OpenAI, any OpenAI-compatible EU endpoint or Claude replace them with one environment variable. Jobs run from a Postgres queue with SKIP LOCKED.
Trade-off
Offline quality is honest but limited: it misses synonyms and cross-lingual questions, and the eval report lists those misses.
Lives in: ADR-005 · ADR-006 · api/app/llm

Mobile and dark.

The same product on a phone, and after hours.
Ask screen on mobile
Contract register on mobile
Navigation menu on mobile
localhost:8080/
Dashboard in dark mode
Dashboard in dark mode
localhost:8080/agent
Agent in dark mode
Agent run in dark mode