The model API is the easy 5%. The rest is plumbing.

Published 3 min read

Calling the model API is often the easy 5% of RAG. The harder part is whether extraction and chunking kept the author's meaning, or fed the model a pile of layout noise.

Setting up a client, sending a prompt, and streaming tokens can feel finished quickly. The real challenge tends to show up when someone uploads a messy company PDF with dual-column layouts, nested financial tables, or scanned pages. Most tutorials skip that part and start from clean English dumps with a simple "split by 500 tokens" recipe.

When the answers go wrong, check the parser first

When production RAG fails, teams often reach for a new model, a bigger context window, or a carefully worded system prompt about tables. Those can help later, but they do not fix corrupted context.

We hit this on real Persian enterprise PDFs. The popular path was something like PyPDF or a LangChain PDFLoader, then chunking by token count. Table rows turned into meaningless fragments, and words broke across ZWNJ (نیم‌فاصله). The embedding store indexed that noise, and the model still answered with full confidence, inventing numbers that looked official. The model was not being clever or stupid. It was reading bad input.

What a flattened budget table looks like

On UGAP, a Persian organizational report had a multi-column budget table. A naive loader flattened it into one text stream, so column headers ended up something like twenty lines away from their values, mixed with prose.

Retrieval still scored high because the keywords matched. The model still had no honest way to know which number belonged to which department. A bigger context window does not fix that on its own, and neither does a cleverer prompt. Layout-aware chunking and Persian glyph normalization do.

I found that part frustrating at first, especially watching people spend days trying to prompt-engineer around answers that were already grounded in bad chunks. It got much better once we dropped the generic loader and wrote a short heuristic pre-processor that rebuilt the tables properly. Not glamorous. Maybe fifty lines. It worked.

Keep the stack simple until it hurts

You usually cannot prompt-engineer your way out of corrupted embeddings. A few practical habits help:

  1. Read raw chunks before you vectorize. If a human cannot answer from the chunk, fix the parser before you blame the model.
  2. Treat tables as structured objects. Do not dump them into a text splitter with no row-level schema.
  3. Normalize language quirks first. In Persian, Unicode and نیم‌فاصله handling is part of the job, not an edge case.

Win the parsing pipe first. The model call is the easy part on top.

Related: WTF is RAG anyway?

چرا صدا زدن مدل فقط ۵ درصد RAG است؟ | Erfan Shafiee Moghaddam