Frequently Asked Questions About Intellipaat Agentic AI Builder Framework

21 answers covering everything from basics to advanced usage.

// Basics

What does 'unstructured in, unstructured out' actually mean?

It means every generative AI system takes input in an unstructured format and produces output in an unstructured format — text, image, audio, or video. It never outputs probabilities, class labels, or regression numbers. If your desired output is a structured value like a fraud score or category, that's a classic machine learning problem, not a generative one. Structured output consistency remains one of the hardest problems in the industry.

What is the difference between the pre-training and adaptation phases of an LLM?

Pre-training is the first phase, where a baseline model learns patterns, word order, token probabilities, and general world knowledge from massive, mostly unstructured data. Adaptation is the second phase, where that baseline is fine-tuned on smaller, specific datasets to specialize its output — for Q&A, image generation, or code. Understanding both prevents over-relying on a base model for specialized tasks it was never adapted for.

Why is token count not the same as word count?

Because LLMs use tokenization strategies like Byte Pair Encoding (BPE), which splits words into subword units based on frequency patterns. A single word like 'ecstatic' may become multiple tokens. This is why you can't estimate cost or context usage by counting words alone — you must count tokens, and both input and output tokens are billed, with output priced higher.

What is a ReAct agent and why is it preferred over PAL?

A ReAct (Reason + Acting) agent first reasons through a problem internally, then acts based on that reasoning. It's the industry-standard architecture implemented by LangGraph, Crew.AI, n8n, and most production frameworks. PAL agents are not used in production. When choosing an agent pattern, default to ReAct — it's proven, widely supported, and the assumption behind nearly every serious agentic framework.

// How To

How do I design a context window budget for my application?

Calculate expected input token size (prompt plus conversation history) and expected output token size, then confirm they fit together within the model's context window. Set a max_token or max_length limit in your app. For production, plan what happens as you approach the limit — truncation strategy, conversation summarization, or a conversation reset — so you never hit silent truncation mid-chat.

How do I build and test for hallucination before deploying?

Test the system with questions where you know the correct answer, then with ambiguous or trick questions. Specifically stress-test timezone conversions, arithmetic, and domain-specific facts — common hallucination triggers. Add answer-validation steps in the agent workflow wherever correctness is critical. Treat hallucination as a structural property, not an edge case, and never push a production system live without this testing.

How do I check an LLM's specifications before choosing it?

For each candidate model, record: input modalities accepted, output modalities produced, context window size in tokens, knowledge cutoff date, input token cost per million, and output token cost per million. Verify these on the provider's pricing page — for example, Google AI Studio for Gemini. Remember output token cost is always higher than input. This record drives both your architecture and your cost projections.

How do I set up guardrails for a RAG application?

Explicitly answer three questions before deployment: What topics should the bot refuse to answer? What data fields are off-limits (PII, phone numbers, internal salaries)? Who is permitted to query what? Then program those constraints into your retrieval and response layers. Guardrails aren't self-imposed by the LLM — a human architect must build them. A fintech bot answering pizza-recipe questions is a real, embarrassing failure mode.

// Troubleshooting

My agent keeps giving outdated answers — what's wrong?

It's likely hitting the LLM's knowledge cutoff — the date after which the model has no training data. Base models can't answer questions about events after this date without external retrieval. Fix it by adding a retrieval mechanism: RAG against an internal database for proprietary data, or internet search for current events. Always check your model's cutoff and design retrieval accordingly.

Why is my conversation getting truncated or breaking mid-chat?

You're likely exceeding the context window by not tracking input plus output tokens together. The window is a shared budget: prompt, conversation history, and model response all count. When it overflows, the system silently truncates. Fix it by budgeting tokens end-to-end, setting a max_token limit, and adding a truncation or summarization strategy for long conversations.

Our GPU costs are killing the project — did we choose wrong?

Very possibly. Locally deployed LLMs cost roughly 32x more than API-based ones due to GPU and infrastructure overhead, and this has killed startup plans mid-build. Unless you're in a regulated industry with strict data sovereignty rules, or your scale economics definitively favor self-hosting, switch to API. Re-run the cost math against actual business value before committing further budget.

The bot answers off-topic or sensitive questions — how do I stop it?

That's a missing-guardrails problem. A RAG or agentic app without explicit constraints will answer anything, including queries that expose PII or breach compliance. Scope retrieval so sensitive fields (phone numbers, salaries, personal emails) are excluded, program topic refusals for off-domain queries, and restrict who can query what. The LLM won't enforce any of this on its own.

// Comparisons

How does agentic AI compare to a traditional rule-based automation?

Rule-based automation follows fixed, predefined logic and is cheaper and more predictable for well-structured tasks. Agentic AI reasons, breaks goals into steps, makes decisions, and uses tools dynamically — ideal for unstructured, multi-step problems. The framework's first principle is to prefer the simpler system: if rules or classic ML solve the problem, use them. Reserve agentic AI for genuinely open-ended, autonomous workflows.

Should I use LangGraph, Crew.AI, or n8n?

Use LangGraph (with LangChain) for maximum customizability and any graph architecture — it's the recommended starting and production framework. Choose Crew.AI for multi-agent orchestration when that's your primary need. Pick n8n for fast no-code or low-code prototyping. All implement ReAct, so the choice is about flexibility versus speed. Avoid unproven protocols like A2A until they demonstrate production stability.

How does RAG compare to fine-tuning for adding knowledge?

RAG retrieves external context at query time, so it's ideal for time-sensitive or proprietary data that changes often, and it keeps data outside the model. Fine-tuning bakes specialized behavior into the model during the adaptation phase, better for consistent task-specific output like document fraud detection. RAG is cheaper to update; fine-tuning is heavier but produces deeply specialized models. Many production systems combine both.

When is MCP better than the A2A protocol?

MCP is better in essentially every current production scenario. MCP (Model Context Protocol) is production-proven for connecting agents to tools, data nodes, and policies. A2A (Agent-to-Agent) is immature and unlikely to survive as a standard. Until A2A proves stability, build async MCP servers for tool access and treat A2A as experimental, not production infrastructure.

// Advanced

How should I approach a document fraud detection build?

First apply the 'not a magical pill' test — check if a simpler system works. If documents are highly variable and unstructured (payslips, statements, IDs), fine-tune the adaptation phase on labeled fraudulent examples across document types. Build multiple validation checkers: metadata analysis, normalization, and chunk-level comparison. Target 95%+ accuracy before production, expect 18+ months of iteration, and budget for purchasing fraudulent-document datasets if none exist internally.

How do I structure a multi-step marketing agent with LangGraph?

Model each task as a node in the graph: research, write, generate image, schedule posts, analyze engagement. Build a ReAct agent that reasons about which node to execute next and acts accordingly. Use API-based LLMs for cost efficiency, define brand guardrails for acceptable content, and specifically hallucination-test the research and fact-generation nodes, since those carry the highest correctness risk.

How do I evaluate cost, energy, and ROI before scaling?

Before scaling, calculate monthly API token cost at projected usage, infrastructure cost if self-hosted, and the business value generated. Every inference burns data-center energy, so generative AI isn't cost-effective at scale without deliberate management. Build cost monitoring into production from day one. If the math doesn't work, a simpler non-generative system may still be the right answer — and that's a legitimate outcome.

When should a regulated enterprise accept the 32x local cost premium?

Accept it only when data sovereignty is non-negotiable — banking, finance, or healthcare use cases where data legally cannot leave the organization's infrastructure. In those cases, the 32x cost of GPU deployment, maintenance, and compliance is the price of admission. Everywhere else, apply the API-first rule. Document the compliance driver explicitly so leadership understands why costs are far higher than an API build.

How do I decide whether a system is generative or agentic early on?

If the system only needs to respond to a single prompt with content, it's a generative AI application. If it needs to set goals, break tasks into steps, make decisions, use tools, or execute multi-step workflows autonomously, it's agentic. Don't conflate them — they differ in architecture, cost, and failure modes. Classifying correctly at step two prevents building the wrong system entirely.