Frequently Asked Questions About Cloud Guru LLM Engineering Production Framework
23 answers covering everything from basics to advanced usage.
// Basics
What does 'the model proposes, your code disposes' actually mean?
It means the model never runs anything — it only produces text describing what it wants done. Every guardrail (approval gates, retries, spending caps, audit logs) lives in your loop, not the model. This reframes reliability as an engineering problem you control rather than a model property you hope for. Design your harness to execute, validate, and refuse based on what the model outputs.
Why is next token prediction the most important intuition to internalise?
Because it explains the number-one failure mode: the model optimises for text that looks right, not text that is right. Every model predicts the most likely next token, appends it, and repeats. Fluency and correctness are different axes. Once you accept this, you stop expecting the model to 'know' facts and start engineering retrieval, validation, and grounding around it.
What is the demo-to-production gap and why do projects die there?
A proof of concept works because you were in the loop watching every output and quietly retrying bad ones by hand. Production means the model runs unattended at scale in front of real users and real money — everything you did implicitly must become a system: retries, validation, guardrails, observability, cost controls, and evals. Most LLM projects die because teams never build that system.
Do all frontier models accept a temperature parameter?
No — the newest frontier models may not accept a temperature parameter at all, exposing an 'effort' or thinking-budget setting instead. Always check the docs for the model you are actually calling before assuming sampling controls apply. Reasoning models are controlled by effort settings and are genuinely better on hard verifiable problems but wasteful and slow on easy ones — route accordingly.
// How To
How do I set up a production-grade generation call?
Read the API key from the environment (never hardcode it), pass model, max_tokens, and a messages list with explicit roles. Always check stop_reason after every call — 'end_turn' means natural completion, 'max_tokens' means the response is truncated and must not be parsed blindly. Log token usage because token counts are your bill. This is step one of the framework's foundation layer.
How do I implement the ReAct pattern for an agent?
Force the model to write a short Thought before each Action (a tool call), then read the Observation (the tool result) before the next Thought. Your code executes the tool_use block and feeds a tool_result back, then the loop runs again. This grounds every step in what actually came back from the world rather than a plan committed to before any data existed. Always cap iterations.
How do I add reflection to improve output quality?
After producing a draft, make a second call whose job is to find flaws — logical gaps, unmet constraints, hallucinated citations — then feed the draft plus critique back for revision. Cap revisions at 2–3 passes because diminishing returns appear fast. The strongest signal comes from external checks (run the test, hit the type-checker), because a model blind to an error in the draft is often blind to it in the critique too.
How do I design memory for a long-running agent?
Use four layers: a scratch pad for intermediate results within one run, session memory as the running message history, and long-term memory written to a durable store (file, database, or vector store) retrieved at each step — which is retrieval in disguise. Decide what to write, when to fetch, and how much. In multi-user systems, scope memory per user, and never let the model write secrets into persistent files.
// Troubleshooting
My RAG answer is wrong — where do I look first?
Suspect retrieval before generation, because most RAG bugs are retrieval bugs. Check whether retrieval even fetched the right document. Common causes: chunking a PDF by fixed character length that slices tables and sentences in half, or embedding your query with a different model than your documents (which silently wrecks results). Instrument observability so you can see the exact context the model was handed for each bad answer.
My structured output validation keeps failing intermittently — what's happening?
Check stop_reason first — if it's 'max_tokens' the JSON is truncated and will never parse, so raise your max_tokens ceiling. If you're streaming, you may be parsing the object token-by-token, which breaks on partial objects; wait for the complete object or use a partial-input parser. And remember the model can emit well-formed JSON that is semantically wrong, so always validate against your schema and retry with the error.
My multi-turn agent seems to forget everything between requests — why?
Because the API is stateless — the model remembers nothing between calls. If you want multi-turn conversation, you must resend the entire message history on every single request. Nothing is carried forward unless you deliberately append it. For persistence across sessions you need long-term memory written to a durable store and retrieved at each step.
My model buries key information and gets facts in long contexts wrong — is that normal?
Yes — it's the lost-in-the-middle effect. Models attend well to the start and end of long contexts but reliably miss information buried in the middle. Related is context rot: as you pack more tokens in, attention spreads thin and the model gets worse at using any single fact, even one it can technically see. More context is not automatically better; position and relevance matter.
Why shouldn't I write a test that asserts exact output from a live model?
Because even temperature zero is not a guarantee of identical output across batches or model updates — it is stable, not deterministic. Asserting byte-for-byte equality on a live model produces brittle tests that break for reasons unrelated to your change. Use an eval set with LLM-as-Judge or structural/semantic checks instead of exact string matching.
// Comparisons
How does hybrid search compare to pure vector search?
Pure dense vector search catches paraphrase and synonyms but can miss exact keywords, product codes, and error strings. Hybrid search runs dense embeddings and BM25 lexical search simultaneously, then fuses the two ranked lists with Reciprocal Rank Fusion — which combines by rank position rather than raw scores, side-stepping the problem that the two systems produce scores on totally different scales. Hybrid is the production default for serious RAG.
How does HNSW compare to IVFPQ for a vector index?
HNSW gives the best recall-latency tradeoff when RAM allows, but it is memory-hungry at hundreds of millions of vectors. IVFPQ searches only the few clusters nearest the query and compresses vectors to slash memory at some accuracy cost. Choose HNSW when RAM is available and recall matters; reach for IVFPQ when the dataset is huge and RAM is the binding constraint.
How does a single well-designed agent compare to a multi-agent system?
Start with a single agent and the right tool surface. Adding multiple agents too early multiplies token cost, latency, and failure modes. Add agents only when you can point to concrete parallelism or context isolation a single loop cannot provide. AutoGen-style debate can improve answers but requires firm termination rules and a turn budget or the agents loop forever.
How does a public benchmark score compare to evaluating on my own data?
A leaderboard number is not a grade for your app. A model that tops SWEBench may be mediocre at your retrieval-heavy support bot. Pick benchmarks whose task shape matches your system — SWEBench for coding agents, TauBench for tool-using customer-service agents — but always validate on your own frozen golden eval set of real inputs with known good answers.
// Advanced
How do I build a three-layer evaluation system?
Layer 1: an offline golden eval set of real inputs with known answers, run on every prompt/model/index change — treat a score drop as a bug. Layer 2: LLM-as-Judge grading open-ended outputs against an explicit rubric. Layer 3: regression testing where every fixed failure joins the set forever. Add red teaming as an adversarial suite covering jailbreaks, injection payloads, and system-prompt extraction, automated on a schedule.
How should I instrument observability for an LLM system?
Trace every request as a tree of spans — top-level request, retrieval calls, model calls, tool invocations — with timing, token counts, and cost at every node. Log actual prompts and completions, because for a non-deterministic model the input and output are the bug report. Apply redaction and retention policies, use OpenTelemetry GenAI semantic conventions so traces sit alongside your other services, pin a spec version, and sample smartly at scale.
How do I safely run code that an LLM generates?
Never run generated code in your own process — a single bad command can wipe files or exfiltrate secrets. Sandbox it in an isolated container or microVM, cap CPU, memory, disk, and wall-clock time, and default to no network access. This turns code execution from a catastrophic risk into a bounded one. It is a non-negotiable guardrail for any coding agent.
What is MCP and why would I build an MCP server?
MCP (Model Context Protocol) is an open JSON-RPC standard — the USB-C of model tooling. Build a server once (tools are verbs the model can invoke, resources are nouns it can read) and every MCP-aware client can use it without a rewrite. Keep each tool narrow, name it clearly, document arguments thoroughly (the description is the entire interface the model sees), scope permissions, and validate inputs.
What is the Clean Core principle for enterprise LLM integration?
In SAP/enterprise contexts, you do not hack the standard system or bury AI calls inside frozen code. You expose stable released APIs and build against those, so when SAP upgrades the system underneath, your integration keeps working. Default to the side-by-side pattern: build the agent as a separate app on BTP that reaches into SAP through OData and RAP. The LLM drafts; deterministic SAP logic commits.
When should I use the batch API instead of synchronous calls?
Use the batch API for throughput-tolerant workloads — overnight document processing, bulk embeddings, eval generation — at roughly half the normal price. Submit the job, poll for completion, then retrieve results within a turnaround window of up to 24 hours. Results come back in any order, so key each request with your own identifier and match on output, never by position.