Frequently Asked Questions About Intellipaat Agentic AI Systems Builder
21 answers covering everything from basics to advanced usage.
// Basics
What does 'unstructured in, unstructured out' mean for generative AI?
It means generative AI takes unstructured input — text, image, audio, video, or code — and produces unstructured output in those same four modalities. It never outputs probabilities, class labels, or regression numbers. If someone claims an LLM directly outputs a probability score or a class label, that violates the fundamentals; those are outputs of traditional ML systems, not generative models.
What is a token and how does it relate to cost?
A token is the fundamental unit of LLM input and output, roughly 4 characters of English text. Words or sub-word fragments are determined by Byte Pair Encoding (BPE), so rare words split into more tokens. Both input and output tokens are charged and count toward the context window. Tracking tokens per session is how you control production cost.
What are the two phases of an LLM's lifecycle?
Every LLM has a pre-training phase and an adaptation phase. Pre-training teaches the baseline model contextual relationships, token probabilities, and general patterns from billions to trillions of mostly unstructured tokens. Adaptation (fine-tuning) specialises that model on smaller task-specific labelled data for a function like Q&A, image generation, or fraud detection. Know which phase your problem lives in before designing.
What is a knowledge cutoff and why does it break my system silently?
A knowledge cutoff is the date after which an LLM has no training data. It breaks systems silently because the model won't say 'I don't know recent events' — it will confidently answer using stale or hallucinated information. If your use-case needs current knowledge, you must augment the model with RAG or real-time tool access, and always check the cutoff before selecting a model.
// How To
How do I qualify whether generative AI is even the right tool?
Ask whether a simpler machine learning system could solve the problem for less cost and energy. Generative AI is only appropriate when the use-case genuinely requires unstructured input or output, reasoning, or language understanding. If a classification or regression model handles it, use that instead. Defaulting to generative AI for everything is an engineering mistake, not innovation.
How do I configure an LLM correctly for production?
For any model you choose, document its four key parameters: input token price per 1M tokens, output token price per 1M tokens, context window size, and knowledge cutoff date. Then set a max_token or max_length limit in every production call to control both cost and behaviour. These four numbers determine viability at scale, not just accuracy.
How do I build a multi-agent content pipeline?
Build it as a LangGraph multi-node workflow where each node is a ReAct agent with access to relevant tools. For a content pipeline: Node 1 (research agent with web search) → Node 2 (script writer) → Node 3 (validator that checks against source material) → Node 4 (scheduler with calendar API). Use API deployment since there's no sensitive data, and cost typically stays under $4/month.
How do I add tool integrations to a LangGraph agent?
Register any external tool the agent needs — search, database, or API — as a tool in LangGraph. For structured communication between AI systems and external nodes, implement Model Context Protocol (MCP). If you need a custom node, build your own async MCP server. Avoid Agent-to-Agent (A2A) protocol since it's a passing trend, not a production standard.
// Troubleshooting
My RAG system is returning off-topic answers and PII — what went wrong?
You built RAG without guardrails, so it answers any question and retrieves any data, including phone numbers and emails. Fix it by explicitly defining accessible data scopes, in-scope query types, and blocked PII fields — then implementing those as explicit filters, not just prompt instructions. The AI will never self-restrict; data scope and access permissions are a human architecture decision.
My conversation started losing earlier context — is this a bug?
No, you've likely exceeded the context window. The context window is the per-conversation token limit combining input prompt, conversation history, and output. Exceeding it causes the oldest tokens to be silently truncated rather than throwing an error. Track both input and output tokens per session, and set limits so long conversations don't quietly drop critical earlier context.
The LLM keeps giving confidently wrong financial figures — how do I fix it?
That's hallucination on high-stakes numerical output, and it must be handled architecturally. Require that all numerical figures come directly from retrieved source records, not generated by the LLM. Add retrieval verification steps, source-citation requirements in the prompt, and human-in-the-loop checkpoints for borderline or high-value cases. Never assume a confident answer is correct just because it sounds authoritative.
My local LLM deployment is blowing the budget — did I make the wrong call?
Possibly — local deployment costs roughly 32x more than an equivalent API due to GPU, infrastructure, and maintenance. Local deployment is only justified for regulated industries or hard data-sovereignty requirements. If you're a startup or mid-market company without strict compliance mandates, migrate to a pay-as-you-go API. The 32x multiplier makes local the exception, not the default.
// Comparisons
How does RAG compare to fine-tuning for adding domain knowledge?
RAG retrieves external knowledge at query time without changing the model, making it cheaper, faster to update, and ideal for proprietary data or recent events. Fine-tuning bakes specialised behaviour into the model weights and is expensive and time-consuming. Try RAG plus prompt engineering first; only fine-tune when those cannot close the accuracy gap on a highly specialised task like fraud detection.
How do ReAct agents compare to PAL agents?
ReAct (Reason + Act) agents reason internally about a task and then act based on that reasoning in a loop — this is what LangGraph, CrewAI, and NIM implement, making it the industry standard. PAL agents are not used in production. For any real system, build ReAct agents; treating PAL as a viable production pattern is a common mistake.
Should I use code-generation tools like Cursor or Lovable to build production AI systems?
Not for security-critical or production-grade software. Tools like Cursor, Lovable, Replit, and Base44 accelerate prototyping but do not produce secure, production-ready code. For agentic systems handling sensitive data, financial figures, or compliance requirements, engineer the code deliberately using LangGraph and LangChain rather than relying on generated output you haven't fully audited.
How does agentic AI compare to a simple ML classification model for cost?
A simple ML classification model is far cheaper and more energy-efficient than generative or agentic AI for tasks it can handle — like structured prediction with class labels. Generative AI is energy-intensive and expensive. The methodology explicitly starts by asking if a simpler ML system works; recommending agentic AI when classification suffices is an engineering error, not sophistication.
// Advanced
How do I decide between a direct LLM call and a full RAG pipeline?
Use a direct LLM call when neither of two RAG-triggering conditions holds: the knowledge cutoff doesn't matter for your use-case, and you don't need to answer questions about internal or proprietary data. If either condition is true — recent events or internal documents — you need RAG. Adding RAG when a direct call suffices just increases complexity and cost.
How do I structure LLM output into a consistent, machine-readable format?
This is where the real engineering value lies — coercing unstructured LLM output into a consistent structured format specific to your system's requirement. It's non-trivial and must be engineered deliberately using output schemas, structured-output modes, validation, and retry logic. Don't assume the model will reliably return valid JSON or a fixed schema from prompting alone; build enforcement around it.
When is fine-tuning genuinely worth the cost?
Fine-tuning is worth it only when API plus RAG plus prompt engineering cannot reach the required accuracy on a highly specialised task — for example, detecting fraudulent bank statements, salary slips, and PAN cards where a general model won't hit 98%+ accuracy. It requires large labelled domain corpora and is expensive and slow, so it's a last resort after exhausting cheaper options.
How do I monitor and control token cost in a production agentic system?
Track input plus output tokens per session, set per-session and daily limits, and use pay-as-you-go pricing where token consumption is unpredictable. Cost, not just accuracy, determines whether a system is viable at scale. Combine token limits with max_token settings on each call so a runaway agent loop can't quietly generate thousands of dollars in charges.
How do I handle a use-case that mixes classification and generation, like fraud detection?
Recognise the primary pattern first. Document fraud detection is fundamentally a classification and anomaly-detection workflow enhanced by generative AI — not a RAG problem. For a regulated bank, deploy locally for data sovereignty, fine-tune on labelled genuine and fraudulent documents across all document types, and add human-review checkpoints for borderline cases since a false positive harms real customers.