TelediRAG IN PRODUCTIONEspañol

Your RAG is not failing on the model. It fails on ingestion.

We build custom RAG systems for organisations that need citable answers rather than plausible approximations, and we measure them against a question set with known answers.

0Answers served without a retrievable source

02 / Why loading the PDFs into a vector database does not work

Why loading the PDFs into a vector database does not work


  • 01

    Blind chunking

    Splitting a document every 512 tokens cuts tables in half and separates a clause from the condition that voids it. The retrieved fragment is syntactically clean and semantically wrong, which is the hardest kind of error to notice.


  • 02

    Structure thrown away

    A PDF is not flat text. It has hierarchy, tables, footnotes and annexes. Losing that structure during ingestion loses the context that makes the number on the page mean something.


  • 03

    No reranking

    Vector similarity retrieves what looks alike, not what is correct. Without a second cross-encoder pass over the candidates, the genuinely useful passage tends to sit in position nine and never reaches the model.


  • 04

    No evaluation set

    If there is no set of real questions with known answers, nobody can claim the system works. It seems to answer well is a feeling, not a measurement, and it does not survive the first disagreement.


  • 05

    No traceability

    An answer that cannot point to the document, page and paragraph it came from is not auditable, and therefore not usable in any process that has consequences attached to it.

03 / How we build it

How we build it

  1. 01

    ingestion

    Structure-preserving parsing: tables, hierarchy and metadata kept intact

  2. 02

    chunking

    Split by semantic unit rather than by a fixed token count

  3. 03

    indexing

    Hybrid index: dense for meaning, lexical for codes and reference numbers

  4. 04

    retrieval

    Hybrid search filtered by metadata and by the user's own permissions

    38 ms

  5. 05

    reranking

    Cross-encoder reordering over the retrieved candidate set

    120 ms

  6. 06

    generation

    Answer constrained to the retrieved context, with a citation required

    684 ms

  7. 07

    evaluation

    Regression set run on every configuration change, before it ships

04 / Integrations

Integrations


  • SharePoint

  • Google Drive

  • Confluence

  • Amazon S3

  • PostgreSQL

  • Snowflake

  • BigQuery

  • Salesforce

  • SAP

  • Dynamics 365

  • Notion

05 / Guarantees

Guarantees


842 ms
p50 latencyFull query, generation included
1.4 s
p95 latency
100 %
Answers carrying a citationWith no retrieved source, the system says it does not know
Continuous
EvaluationRegression set executed on every deployment

06 / Compliance

Compliance


  • Data hosted exclusively inside the European Union

    GDPR, chapter V


  • Access control propagated from your source system into the index

    GDPR art. 5(1)(c), data minimisation


  • Query log and retrieved sources retained for audit

    EU AI Act, traceability obligations, applicable from 2 August 2026


  • No model retraining on your documents

    Data processing agreement


  • Every answer traceable to the document, page and paragraph it came from, in a form an inspection can follow

    EU AI Act art. 99(4), penalties of up to EUR 15M or 3 % of global annual turnover

07 / Process

Process


  1. 1-2

    Corpus audit

    Inventory of sources, formats and permissions, and identification of contradictory or superseded documents. Those contradictions are the single most common cause of wrong answers, and they exist before any model is involved.


  2. 2-3

    Evaluation set

    We build a set of real questions with known answers together with your team. Without this piece there is no objective way to tell whether a change made the system better or worse.


  3. 3-6

    Ingestion and indexing

    Parsing, chunking and the hybrid index, iterated against the evaluation set until it clears the threshold agreed at the start rather than the one that happens to be reachable.


  4. 6-9

    Integration and deployment

    Connection to your systems, permission propagation, observability and production rollout behind the access controls your users already have.


  5. ongoing

    Operation

    Periodic re-evaluation, incremental reindexing and tuning as the corpus changes underneath the system.

08 / Frequently asked questions

Frequently asked questions

How long does a full RAG project take?
Six to nine weeks to production for a corpus of moderate complexity. The variable that matters is not document volume but how many distinct formats and internal contradictions the corpus contains.
What happens when the system cannot find the answer?
It says it does not have enough information. We configure retrieval so that the absence of a retrieved source blocks generation, because an invented answer costs far more than a non-answer.
Can it run over confidential documentation?
Yes. Access control from your source system propagates into the index, so each user only ever retrieves fragments of documents they were already entitled to read. Data stays hosted inside the European Union.
Which models do you use?
We choose per case, based on the language, the domain and data residency requirements, and we design the system so the model is replaceable. The architecture is not wired to one vendor's SDK.
How do you measure that the RAG actually works?
With an evaluation set built with your team: real questions with known answers, run on every configuration change. We measure retrieval precision, answer correctness and the share of answers with a valid citation.
Can you improve a RAG system we already run?
Yes, and it is a large part of our work. We start by building the evaluation set that almost never exists, measure the current system against it, then prioritise fixes by their effect on that measurement.
What happens when the documents change?
Reindexing is incremental: only what changed is reprocessed. For corpora with frequent updates we set up automatic synchronisation from the source system, so the index does not silently drift out of date.

09 / Related services

Related services

10

Tell us which process to fix

Describe the process and the systems behind it. You get back a technical proposal — architecture, timeline and acceptance criteria — not a service catalogue.

Email hola@teledi.ai