An AI assistant that reads a company's contracts and invoices, cites the exact sentence, and never acts without approval.

Overview.
Small and mid-sized companies
A manufacturer with a few dozen supplier contracts, leases, software licences and invoices in Italian, English and sometimes Arabic, and no legal team. The demo company is a fictional shoe maker in Civitanova Marche.
Deadlines hide inside PDFs
Auto-renewal clauses with 60 or 90 days' notice slip by, invoices get paid late, and answering "what does our lease say about…" means opening files one by one. Public chatbots are off the table for company data.
Grounded answers, deadlines and a careful agent
Upload the documents once. Sanad extracts the terms, builds the contract register, calendar and cash-flow projection, answers questions with verified citations, and drafts reminders and letters for approval.
Walkthrough.
One company, one key, one isolated workspace
Each customer company is a tenant with its own API key. Signing in opens a workspace that only ever sees that company's documents, and the demo workspace is a fictional shoe maker in the Marche region.

- 1Keys stored as hashesOnly the SHA-256 of each API key is stored. A key maps to exactly one tenant, and per-key rate limits apply to every request.
- 2Isolation in the databaseEvery request runs in a transaction that switches to a non-privileged role and sets the tenant id. Without that setting, queries return nothing.
- 3Realistic demo dataFour contracts, two invoices and two policies in English, Italian and Arabic, including a lease in .docx and a policy in plain text.
What the company owes, and what needs a decision
The dashboard turns eight documents into numbers an owner cares about: monthly commitments, the next deadlines, projected outgoing cash and how many agent proposals are waiting.

- 1Commitments, not files€16,950 a month in the demo, computed from extracted fees and billing frequency, with a 12-month total of €219K.
- 2Deadlines with countdownsInvoice due dates and the last day to give notice on each contract, each with days left and a link to the source.
- 3Approvals surfacedPending agent proposals are counted on the dashboard and in the sidebar badge, so nothing waits unnoticed.
Answers with the exact sentence and page, or none at all
Ask a question in English, Italian or Arabic. Sanad answers only from the company's own files and shows where each quote came from. If it can't support an answer, it says so.


- 1Verified quotesEach quote is matched back against the retrieved text before it reaches the screen. Citations that don't match are dropped and confidence goes down.
- 2Three languagesThe same pipeline answers in English, Italian and Arabic, including an Arabic question about an Arabic policy.
- 3Guardrails that explain themselves"Ignore all previous instructions" is blocked before any model call, with the matched patterns and score shown instead of a silent failure.
Drop in PDFs, Word files or text, and get structured terms back
Upload contracts, invoices and policies in any of four formats. A background worker parses, chunks, embeds and classifies each file, then extracts the terms that matter.


- 1Type detected automaticallyContract, invoice, policy or other, with the language detected per file. Duplicates are caught by content hash.
- 2Terms, not summariesParties, term, auto-renewal, notice period, fee and frequency, payment terms, penalties and governing law, each tied to the pages it came from.
- 3Implausible values rejectedA notice period of 900 days or a negative fee never reaches the register. Re-running extraction is one click.
A register that knows the last day to cancel
The contract register shows every agreement with its renewal rule and the date by which notice must be sent. Invoices sit next to it with due dates and totals.


- 1Notice deadline, computedEnd of term minus the notice period, adjusted for renewal cycles. It's the date the product exists for.
- 2Value and payment termsFee, billing frequency and payment days per contract, so cash flow can be projected from them.
- 3Evidence one click awayEvery row links back to the source document and pages the value was extracted from.
Deadlines on a calendar, commitments on a chart
Notice deadlines, invoice due dates and manual reminders share one calendar. The cash-flow view projects twelve months of committed outgoing payments from contract fees and open invoices.


- 1Reminders you controlCreate a reminder by hand or approve one the agent proposed. Approved reminders are e-mailed when SMTP is configured, or logged.
- 2Projection from termsNext 12 months, monthly average, peak month and the amount that depends on renewals you could still cancel.
- 3Contracts vs invoicesEach month splits fixed contract fees from open invoices, so the owner sees which part is negotiable.
It investigates freely, and asks before it acts
Give Sanad a goal such as "find the most urgent renewal and draft a cancellation letter". A LangGraph agent reads the register and deadlines, then pauses with proposals that a person approves or rejects.


- 1Read tools run, action tools waitListing contracts and deadlines runs immediately. Creating a reminder or drafting a letter becomes a proposal and the graph stops.
- 2Approval can come days laterRuns are checkpointed in Postgres, so a decision can arrive from another process after a restart. Undecided proposals count as rejected.
- 3A full traceEvery step, tool call, result, proposal and decision is stored and audit-logged. Letters are saved, never sent.
What every answer costs, per company
Tokens, cost and latency are metered for every model call and grouped by tenant, operation and model. An audit log records every human decision and executed action.

- 1Cost per operationEmbedding, classification, extraction, answering and agent steps are metered separately, so the expensive part is obvious.
- 2Provider-neutral numbersThe same view works offline, with OpenAI or an EU OpenAI-compatible endpoint, or with Claude.
- 3An audit trail by defaultApprovals, rejections and executed actions are written as audit events, with who decided and when.
Evaluation.
Retrieval · 35 answerable questions
| Mode | hit@1 | hit@3 | MRR |
|---|---|---|---|
| Vector only | 0.91 | 0.97 | 0.95 |
| Keyword only | 0.94 | 1.00 | 0.97 |
| Hybrid (RRF) | 0.94 | 1.00 | 0.97 |
Answers and extraction
| Metric | Score |
|---|---|
| Answer correctness | 0.94 |
| Citation accuracy | 0.94 |
| Groundedness | 1.00 |
| Abstention on unanswerable | 0.80 |
| Prompt-injection block rate | 1.00 |
| Extraction field accuracy | 1.00 |
Experiment · chunk size and overlap
| Chunk / overlap | hybrid hit@3 | MRR | correctness |
|---|---|---|---|
| 400 / 60 | 1.00 | 0.94 | 0.86 |
| 800 / 120 | 1.00 | 0.97 | 0.94 |
| 1200 / 180 | 1.00 | 0.97 | 0.91 |
What the numbers mean
Scores are from the offline provider, so every run is free and repeatable. Adding contextual chunk headers raised vector-only hit@1 from 0.85 to 0.91 and hybrid hit@3 from 0.97 to 1.00. The golden set covers English, Italian and Arabic, unanswerable questions and injection attempts. The remaining misses are one cross-lingual question and two near-misses, and the 1.00 extraction accuracy is a regression baseline on the sample documents, not a claim about unseen contracts.
Architecture.
Engineering deep dives.
1. Citations are verified, not trusted
- Problem
- Language models invent convincing quotes. In a product about contracts, one fake clause destroys trust.
- Approach
- The model cites numbered sources. Each quote is normalised and matched against the retrieved chunk, exactly or on at least 90% of its tokens to tolerate OCR noise. Unverifiable citations are dropped and confidence is lowered.
- Trade-off
- A correct answer that paraphrases instead of quoting can lose its citation and show lower confidence.
2. Tenant isolation lives in Postgres
- Problem
- A leak between two companies' contracts would end the product. "Remember to add WHERE tenant_id" is not a control.
- Approach
- Every tenant table has a row-level security policy for reads and writes. Each request runs SET LOCAL ROLE to a non-privileged role and sets app.tenant_id for that transaction only. A missing context fails closed and returns nothing. The test suite proves it with raw SELECTs.
- Trade-off
- HNSW scans first and RLS filters after, so a small tenant among many can get fewer than k vector candidates. ef_search is raised per query.
3. Approval only where it matters
- Problem
- An agent that reviews contracts is useful. One that creates reminders or writes to suppliers on its own is a liability.
- Approach
- Tools come in two classes. Read tools run as soon as the model calls them. Action tools never execute when called: they become proposals and the LangGraph graph pauses on interrupt(). State is checkpointed in Postgres.
- Trade-off
- Stopping before every tool call would train people to click approve without reading, so harmless reads run unattended by design.
4. Hybrid search, measured
- Problem
- Business documents mix prose with identifiers that embeddings handle badly: invoice numbers, IBANs, supplier names, clause numbers.
- Approach
- Vector search and Postgres full-text run in one SQL statement and are fused with Reciprocal Rank Fusion. Every chunk is embedded with its document name and title, and each result keeps its vector and keyword rank for debugging.
- Trade-off
- Two indexes to maintain. With offline lexical embeddings, keyword search alone is as strong as hybrid on this corpus.
5. Retrieved text is untrusted input
- Problem
- A malicious instruction can hide inside an uploaded document, not just in the user's question. Company data shouldn't leave in clear text either.
- Approach
- Prompt-injection scanning runs on the question and on every retrieved chunk before it enters the prompt. IBANs, e-mails, phone numbers and Italian tax codes are masked reversibly before any external model call, and restored in the answer.
- Trade-off
- Pattern-based detection needs upkeep as new attack phrasings appear.
6. An offline provider as a first-class citizen
- Problem
- Tests, CI evals and first-run demos shouldn't need an API key or cost money, and some customers will require EU-hosted models.
- Approach
- Every task has a deterministic implementation: lexical embeddings, extractive answers, regex extraction and a scripted agent policy. OpenAI, any OpenAI-compatible EU endpoint or Claude replace them with one environment variable. Jobs run from a Postgres queue with SKIP LOCKED.
- Trade-off
- Offline quality is honest but limited: it misses synonyms and cross-lingual questions, and the eval report lists those misses.
Mobile and dark.





What's next.
Where it stands
Sanad runs end to end with Docker in one command and is tested against a realistic demo company. It has no paying customers yet. The next step is a pilot with a real SME and a run of the same eval suite on real models and real contracts.
Roadmap
OCR for scanned PDFs and table-aware parsing, cross-encoder re-ranking and query rewriting for cross-lingual questions, forwarding invoices to a tenant inbox and importing Italian e-invoices (SDI XML), SSO with per-user roles, and API-key rotation.