[ 05 — AI SYSTEMS ]
Assistants, search and document pipelines built on your own data.
[ 01 — THE PROBLEM ]
What this usually looks like
Before
- A general chatbot knows nothing about your contracts, your stock or your policies.
- The answer exists in a document somewhere and nobody can find which one.
- New staff spend their first month asking colleagues questions the handbook already answers.
- An AI tool gave a confident answer that was wrong, and nobody could tell where it came from.
After
- Your own documents indexed and searchable by meaning, not just by keyword.
- Answers that cite the source document and page, so they can be checked in one click.
- An assistant that knows your policies because it was given them, available where people already work.
- Permissions enforced at retrieval, so nobody is shown a document they are not allowed to read.
[ 02 — WHAT YOU GET ]
Deliverables, not deliverable-sounding words
A document pipeline
Ingestion, parsing, chunking and embedding for your file formats, including the scanned ones.
A search index
Hybrid keyword and semantic search, tuned on real questions from your team rather than on a demo set.
The assistant
In a web interface, or inside Slack, WhatsApp or your existing system — wherever the question is actually asked.
Citations and access control
Every answer traceable to its source, and retrieval filtered by what the person asking is allowed to see.
An evaluation set
A fixed set of questions with known good answers, so a change to the system can be measured rather than guessed at.
[ 03 — HOW IT WORKS ]
The steps, and how long each one takes
- 01
Gather the corpus
3–5 daysWhat documents exist, where they live, what state they are in and who may see them.
- 02
Build the pipeline
1–3 weeksParsing and indexing, including OCR for scans. This is most of the work and most of the quality.
- 03
Tune retrieval
1–2 weeksAgainst real questions from your team. Measured against the evaluation set, not against impressions.
- 04
Ship the interface
1–2 weeksWhere people already work, with citations and a way to report a bad answer.
- 05
Keep it current
ongoingNew and changed documents reindexed automatically, so the assistant does not slowly go stale.
[ 04 — EXAMPLES ]
What people use this for
Policy and handbook assistant
Staff ask in plain language and get the answer with the clause it came from.
Contract and tender search
Find every agreement with a particular clause, obligation or renewal date across years of documents.
Technical support knowledge base
Agents get a drafted answer with sources while the customer is still on the line.
A general chatbot knows nothing about your contracts, your stock or your policies. We build the layer that gives it that knowledge.
Retrieval first
The quality of one of these systems is decided almost entirely by the unglamorous half: parsing documents properly, including the scanned ones, chunking them sensibly, and retrieving the right passage. The model on top matters less than people expect. We spend our time where the quality actually comes from.
Answers you can check
Every answer cites its source. If the system cannot find anything relevant, it says so instead of inventing something plausible — which is the single most important behaviour and the one most demos skip.
[ 05 — STACK ]
What we build it with
Chosen because they are boring, well supported and easy to hire for. Nothing here will be abandoned next year.
- Python
- FastAPI
- Claude
- OpenAI
- pgvector
- PostgreSQL
- Whisper
- LangChain
[ 08 — FAQ ]
AI Systems: questions people ask
Let's talk about ai systems.
Tell us the problem. We will tell you what we would build, what it costs and how long it takes — before you commit to anything.