Your RAG is not failing on the model. It fails on ingestion.
We build custom RAG systems for organisations that need citable answers rather than plausible approximations, and we measure them against a question set with known answers.
02 / Why loading the PDFs into a vector database does not work
Why loading the PDFs into a vector database does not work
- 01
Blind chunking
Splitting a document every 512 tokens cuts tables in half and separates a clause from the condition that voids it. The retrieved fragment is syntactically clean and semantically wrong, which is the hardest kind of error to notice.
- 02
Structure thrown away
A PDF is not flat text. It has hierarchy, tables, footnotes and annexes. Losing that structure during ingestion loses the context that makes the number on the page mean something.
- 03
No reranking
Vector similarity retrieves what looks alike, not what is correct. Without a second cross-encoder pass over the candidates, the genuinely useful passage tends to sit in position nine and never reaches the model.
- 04
No evaluation set
If there is no set of real questions with known answers, nobody can claim the system works. It seems to answer well is a feeling, not a measurement, and it does not survive the first disagreement.
- 05
No traceability
An answer that cannot point to the document, page and paragraph it came from is not auditable, and therefore not usable in any process that has consequences attached to it.
03 / How we build it
How we build it
- 01
ingestion
Structure-preserving parsing: tables, hierarchy and metadata kept intact
- 02
chunking
Split by semantic unit rather than by a fixed token count
- 03
indexing
Hybrid index: dense for meaning, lexical for codes and reference numbers
- 04
retrieval
Hybrid search filtered by metadata and by the user's own permissions
38 ms
- 05
reranking
Cross-encoder reordering over the retrieved candidate set
120 ms
- 06
generation
Answer constrained to the retrieved context, with a citation required
684 ms
- 07
evaluation
Regression set run on every configuration change, before it ships
04 / Integrations
Integrations
SharePoint
Google Drive
Confluence
Amazon S3
PostgreSQL
Snowflake
BigQuery
Salesforce
SAP
Dynamics 365
Notion
05 / Guarantees
Guarantees
- 842 ms
- p50 latencyFull query, generation included
- 1.4 s
- p95 latency
- 100 %
- Answers carrying a citationWith no retrieved source, the system says it does not know
- Continuous
- EvaluationRegression set executed on every deployment
06 / Compliance
Compliance
Data hosted exclusively inside the European Union
GDPR, chapter V
Access control propagated from your source system into the index
GDPR art. 5(1)(c), data minimisation
Query log and retrieved sources retained for audit
EU AI Act, traceability obligations, applicable from 2 August 2026
No model retraining on your documents
Data processing agreement
Every answer traceable to the document, page and paragraph it came from, in a form an inspection can follow
EU AI Act art. 99(4), penalties of up to EUR 15M or 3 % of global annual turnover
07 / Process
Process
- 1-2
Corpus audit
Inventory of sources, formats and permissions, and identification of contradictory or superseded documents. Those contradictions are the single most common cause of wrong answers, and they exist before any model is involved.
- 2-3
Evaluation set
We build a set of real questions with known answers together with your team. Without this piece there is no objective way to tell whether a change made the system better or worse.
- 3-6
Ingestion and indexing
Parsing, chunking and the hybrid index, iterated against the evaluation set until it clears the threshold agreed at the start rather than the one that happens to be reachable.
- 6-9
Integration and deployment
Connection to your systems, permission propagation, observability and production rollout behind the access controls your users already have.
- ongoing
Operation
Periodic re-evaluation, incremental reindexing and tuning as the corpus changes underneath the system.
08 / Frequently asked questions
Frequently asked questions
- How long does a full RAG project take?
- Six to nine weeks to production for a corpus of moderate complexity. The variable that matters is not document volume but how many distinct formats and internal contradictions the corpus contains.
- What happens when the system cannot find the answer?
- It says it does not have enough information. We configure retrieval so that the absence of a retrieved source blocks generation, because an invented answer costs far more than a non-answer.
- Can it run over confidential documentation?
- Yes. Access control from your source system propagates into the index, so each user only ever retrieves fragments of documents they were already entitled to read. Data stays hosted inside the European Union.
- Which models do you use?
- We choose per case, based on the language, the domain and data residency requirements, and we design the system so the model is replaceable. The architecture is not wired to one vendor's SDK.
- How do you measure that the RAG actually works?
- With an evaluation set built with your team: real questions with known answers, run on every configuration change. We measure retrieval precision, answer correctness and the share of answers with a valid citation.
- Can you improve a RAG system we already run?
- Yes, and it is a large part of our work. We start by building the evaluation set that almost never exists, measure the current system against it, then prioritise fixes by their effect on that measurement.
- What happens when the documents change?
- Reindexing is incremental: only what changed is reprocessed. For corpora with frequent updates we set up automatic synchronisation from the source system, so the index does not silently drift out of date.
09 / Related services
Related services
Wiring every agent to every system does not scale.Each direct integration is a promise of permanent maintenance. We build custom MCP servers that centralise how your agents reach your systems, with a tool contract and a schema validated on every call.
An agent with access to everything is a risk, not a gain.An agent with unrestricted access and no audit log is not a productivity tool; it is a security incident with no date on it yet. We build internal agents with role-scoped permissions and a log that survives a security review.
Services
10
Tell us which process to fix
Describe the process and the systems behind it. You get back a technical proposal — architecture, timeline and acceptance criteria — not a service catalogue.
Email hola@teledi.ai