Retrieval-augmented AI — built and run for you

Your content already
has the answers

They're in your website, procedures, spec sheets, contracts and product pages. We build assistants that retrieve the right information and answer from it — with the document, the section and the page attached. For the people evaluating you on the outside, and for your team on the inside.

Your documents are not training data. Nothing you index teaches anyone else's model.
Running in production on a live commercial site
Answers cite the document and the page
Runs in your own cloud where compliance requires it
One engineer who knows your build — not a ticket queue
Grounded
Your docs
Not the model's memory2
Current
Same day
Edit a doc, the answers follow2
Honest
“I don't know”
Rather than a confident guess
Fast
Seconds
Question to sourced answer

It isn't a knowledge problem.
It's a retrieval problem.

Nobody is short of documentation. They're short of a way to get one specific paragraph out of eleven years of it — in the ninety seconds before the customer is back on the phone, or before the visitor closes the tab.

Six systems, one answer

SharePoint, the shared drive, a folder of supplier PDFs, the ticket history, the product pages, and one person's inbox. The answer exists in exactly one of them and nobody remembers which.

Keywords don't match questions

Your search box needs the words that are in the document. People have the words the customer just used. “Can we run it wet?” never matches a heading that reads “Environmental rating, IP69K.”

Someone pays for the gap

Inside, the question goes to your most senior and most interrupted employee, whose knowledge leaves when they do. Outside, the visitor doesn't ask anyone — they close the tab, and you never learn what they wanted.

A general model doesn't know you

A model's knowledge is fixed when its training ends.2 It has never read your warranty policy or your tolerance tables, so when asked it produces something plausible and wrong. Confident and wrong is the expensive failure.

Retrieve first, then answer

We find the passages in your material that actually address the question, and the model is only allowed to answer from those. It reads your documents at the moment of asking, every time.

Show your working

Every answer carries the document, the section and the page it came from. Anyone can verify in one click — the only reason a person will ever trust it enough to use it on a live call, or act on it as a buyer.

One engine,
pointed in two directions

The same retrieval system, the same citations, the same refusal to invent. What changes is who's asking, and what happens at the end of the conversation.

Inside — your team

The question your best person gets asked twice a week

One question box over the SOPs, the spec sheets, the contracts and the manuals. Your staff get the answer and the page it came from, instead of walking over to the one person who knows.

  • Answers cite the document, section and page
  • Scoped to the material you choose to index
  • Unanswered questions routed to the document owner
  • Self-hosted where policy or a regulator requires it

Start with: sales engineering, field service, support, operations — whoever owns the specs. The team that gets interrupted most already knows which questions repeat.

Outside — your buyers

The visitor who won't fill in a contact form

People evaluating something don't ask once — they compare, push back and narrow down. The assistant answers from your published material, and when someone signals real intent, a human is told within minutes.

  • Deep links into live pages and PDF pages
  • Intent detected mid-conversation, not at a form
  • Routed to CRM, Slack or inbox with the full transcript
  • Indexes whitepapers behind a form, without breaking the form

Start with: the pages where evaluation happens — products, specs, comparisons, pricing. That's where the questions are worth money.

Open-book AI

A closed-book model answers from memory and hopes. This one looks the answer up in your documents first, answers only from what it found, and shows you the page. Four steps — and the third is a hard constraint, not a suggestion.

01

Index

Your pages and documents are split into passages and stored so they can be found by meaning rather than by keyword. Text, tables and PDFs, including the scanned ones.

Once, then kept current
02

Retrieve

A question comes in. We pull the handful of passages that genuinely address it — matching on meaning, so the question doesn't have to use the document's vocabulary.

Under a second
03

Ground

Those passages go to the model with strict instructions: answer from this material only. Nothing outside it is available to draw on, so nothing outside it can be invented.

Every single answer
04

Answer

A written answer in plain language, each part linked back to the passage behind it. Where the passages don't cover the question, it says so and points at a human.

With sources attached

Because the knowledge lives outside the model, nothing has to be retrained when your material changes. Correct a procedure this morning, re-index, and this afternoon's answers reflect the correction.2

Down to the paragraph

An answer nobody can check is a rumour. Every claim is tied to the exact paragraph it came from, with a deep link into the live page or the specific page of the PDF — so the reader can verify it, and so you can audit it later.

  • Paragraph-level, not page-level, sourcing
  • Deep links into live pages and PDF pages
  • Full audit trail of every answer and its sources
Answer audit log
Answered · 3 sources
"Wash-down duty rating for the 400 series"
spec-sheet-400.pdf · p.7
maintenance-schedule-r6.docx · §3.2
Redacted · retention 90d

You can see when it's struggling

Every conversation is scored from real signals — thumbs, fallbacks, abandonment, escalations, whether it ended in something useful. Filter to the bad sessions and read exactly what went wrong, instead of guessing whether the thing is working.

  • Good / neutral / bad session labels, with reasons
  • Drill from any chart straight into the transcripts
  • Token and cost usage tracked per conversation
Session Quality
Good · resolved + lead
Neutral · answered, left
Bad · 2 fallbacks, abandoned

Every gap is a signal

The questions your material can't answer aren't just documentation gaps — they're demand. We rank them by how often they're asked and pair each with who asked, so the right person knows what to write and, on a public site, who to call.

  • Ranked unanswered topics, clustered automatically
  • Monthly digest of what's missing and what's rising
  • Which documents get cited — and which never do
Unanswered Questions
"Lead times for the 400 series" 47 asks
"Does it integrate with our ERP?" 31 asks
"Warranty terms outside North America" 28 asks
"Approved substitute parts" 19 asks

It stays yours

This is the first question every serious buyer asks, and it deserves a direct answer rather than a page of assurances. Your documents are indexed into a store you own. They are not used to train anything. Where your policy or your regulator requires it, the whole system runs inside your own infrastructure and nothing leaves it.

  • Your content is never training data
  • Self-hosted deployment where it's required
  • Retention rules and transcript redaction you set
  • No third-party cookies — passes security review
New enquiry · 4 min ago
Ready for a call
Asked about throughput, ERP integration and delivery windows across 6 turns. Landed from an organic search on the 400-series page.
Specification Integration Timeline
4 sources cited · 6-turn session

Two exchanges.
Watch the second answer.

In both cases the assistant refuses the question it can't source, says why, and turns the dead end into a next step. That behaviour is the whole product.

internal assistant · sales engineering
Kate, inside sales
Distributor in Alberta is asking whether the extended warranty still covers wash-down failures. Do we need them to co-sign anything?
Assistant
Yes, wash-down failures stay covered under the extended warranty, on two conditions: the documented wash-down procedure is followed, and seal inspections are logged at the stated intervals. For distributor-sold units the distributor signs the service acknowledgement — the customer does not. Steam above 80 °C voids coverage entirely.
Kate, inside sales
And if the unit was bought before the 2026 policy came in?
Assistant
I only have the 2026 policy indexed — I don't have the superseded version, so I can't tell you what applied before it and I'm not going to guess on a warranty question. Rick in Contracts is listed as the owner of that document set. Want me to send him this thread?
No source — not answered Gap logged: pre-2026 policy Routed to document owner

A system that guesses at the second question is worse than no system, because your team will believe it once and stop believing it forever. Refusing to answer, naming the missing document, and finding the person who owns it is what makes people willing to use it on a live call.

You have three options.
Two of them are fine.

We'd rather you picked the right one than picked us. Here is the honest comparison, including where a cheap off-the-shelf tool beats us outright.

Build in-house Off-the-shelf tool Work with us
Time to something useful 8–14 weeks for a median build — then it is never really finished1 An afternoon Two to three weeks
Up-front cost $15,000–$40,000 for a basic build. The median is $75,000–$120,0001 Nothing — though enterprise platforms want $12,000–$50,000 a year on an annual contract4 A fixed project fee
What it costs to run $150–$1,500 a month in tokens and hosting3, plus 15–25% of the build price every year in maintenance1 $15 a month to start. $150–$1,500 once people actually use it3 One monthly fee. Nothing metered, no usage bill.
Who you need 2–4 developers with retrieval and vector database skills, at $200–$300 an hour blended1 Anyone who can drag a file Nobody. That's the point.
Handles messy PDFs, tables, scans Eventually, if you fund it Poorly, and it won't tell you Yes — it's most of the work
Tuned to your questions Yes, by your team, forever No. Generic retrieval, generic results Yes, by us, against real transcripts
When it gets an answer wrong Your team debugs retrieval You file a ticket and wait We fix it, and tell you why it happened
Best when This capability is core to your product and you're staffing for it Your content is tidy, public and low-stakes Answers must be right, and the documents are a mess

If your documentation is clean, in one place, and nobody gets hurt by a wrong answer, buy an off-the-shelf tool for a few hundred a month and don't call us. We're worth the money when the material is difficult and the cost of being confidently wrong is real.

Start narrow,
and prove it there

Every rollout that fails begins with “let's index everything.” Pick the smallest set of material that causes the most pain, and find out whether this works before you widen it.

Choose the team, or the pages, that get hit most. Inside, that's usually sales engineering, service, or whoever owns the specs. Outside, it's the product and comparison pages where evaluation actually happens. Both already know which questions repeat, which is the fastest tuning signal there is.

Give us the awkward documents, not the tidy ones. A clean wiki proves nothing. Hand over the scanned manuals and the spreadsheet with merged cells — if it works on those, it will work on everything else you own.

Agree what “right” means before we start. We write down twenty real questions and the answers you'd accept. That list is how the build gets judged, and it stops the project ending in vague opinions about whether the AI is any good.

Expect the gap report to sting a little. The list of questions your material can't answer is genuinely useful and mildly embarrassing for everyone. It's often worth more in the first month than the assistant is.

One fixed build fee,
then a year of keeping it right

An assistant that isn't maintained goes stale within a quarter — your material moves and the answers stop matching it. So the build and the year that follows are sold together, and quoted together, after one call.

One — the build
Fixed fee
Live in about three weeks
Quoted once we've seen your material. No hourly billing, no change orders for things we should have anticipated.
Your documents indexed, the retrieval tuned by hand, and the assistant judged against real questions before anyone calls it finished.
  • PDFs, spec sheets, scans and gated documents
  • Citations and strict grounding on every answer
  • Internal deployment, public widget, or both
  • CRM, Slack and calendar handoff
  • Judged against twenty questions you choose
  • Gap report on what your material can't answer
Two — the year
Monthly
Hosting, maintenance and tuning
One monthly fee covering everything it costs to run and everything it takes to keep accurate. Nothing metered, no surprise usage bill.
Someone keeping it right as your material, your questions and the models underneath it all change — which they will, all three, within the year.
  • Hosting and all running costs included
  • Re-indexing as your documents change
  • Monthly tuning against real transcripts
  • Wrong answers fixed, with an explanation of why
  • Model upgrades handled, at no extra cost
  • Monthly report: questions, quality, gaps, cost
Three — the term
12 months
Then month to month
A year is the minimum because tuning compounds — the assistant is measurably better in month nine than month two, and that only happens if someone is still working on it.
Straightforward terms, written down before you sign, with nothing that quietly renews itself while you aren't looking.
  • Twelve months from go-live, then month to month
  • No auto-renewing year you have to escape
  • Half the build on signature, half on go-live
  • Your content, conversations and leads — exportable at any time, including on the way out
  • Leave early and the balance is payable — so tell us in month three, not month eleven

What it costs depends on what you've got

The build fee turns on how much material there is and what state it's in — a tidy documentation set and forty thousand scanned pages are not the same job. Twenty minutes on a call is enough for us to put a fixed number in front of you, and to tell you if the honest answer is that you don't need us.

Not ready to commit to a build? We also run a paid three-week pilot on one content set, credited in full against the build if you go ahead — and yours to keep if you don't.

We take on a small number of builds at a time, because every one is tuned by hand against your material. If that means waiting a few weeks for a slot, we'll say so on the first call.

Bring one awkward question

The fastest way to work out whether this is worth your time is to give us a question your team or your website currently answers badly — and roughly which document holds the answer. That's enough for us to tell you whether this is worth doing.

A first call is twenty minutes. If your material isn't in a state where this would work yet, we'll say so and tell you what would need to change.

We reply personally, usually same day. No sequences, no newsletter.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Sources. 1 SFAI Labs, RAG Development Costs — a development agency's own published tiers: basic $15,000–$40,000 (3–6 weeks, 1–2 developers); moderate $40,000–$120,000 (8–14 weeks, 2–4 developers); complex $120,000–$250,000; enterprise $250,000–$500,000+. Median project “$75,000–$120,000 with 8–14 weeks of development time.” Maintenance runs “15–25% of initial development costs annually,” on blended US agency rates of $200–$300 an hour. Compliance work adds 30–100%, integrations 20–60%, and a compressed timeline 30–50%. 2 indigo.ai, Retrieval-Augmented Generation — an LLM's knowledge is frozen when pre-training ends; RAG updates the knowledge base without changing model parameters, and citing sources is what makes the output checkable. 3 SpendArk, RAG System Cost — monthly running cost of a live RAG system: $150–$400 at 10,000 queries a month, $600–$1,500 at 150,000, rising to $5,000–$15,000 at a million. LLM inference is roughly 68% of that bill; the vector database is $25–$500 and embeddings $5–$50. The same workload costs ~$81 a month on a small model and ~$1,350 on a large one — which is why a subscription price and a real usage bill are not the same number. 4 Context Link, RAG as a Service — managed RAG platforms span $9–$19 a month for small-business tools, $100–$1,500 a month for platform builders, and $12,000–$50,000+ a year for enterprise infrastructure, the last “often requiring annual contracts,” “4–12 weeks of implementation” and “dedicated engineering resources.” Further reading: Forbes Business Council, Unleashing The Power Of Retrieval-Augmented Generation For Medium-Sized Businesses, also worth reading for its warning about vendors who use AI terminology to imply expertise they don't have.