Frequently Asked Questions About Rajeev Kanth Agentic AI Architecture Stack
22 answers covering everything from basics to advanced usage.
// Basics
What are the four core characteristics of agentic AI?
The four core characteristics are autonomy, tool use, memory, and reflection. Autonomy means the agent self-corrects its own errors through loops. Tool use connects the LLM to external APIs to take real-world actions. Memory includes short-term conversation history and long-term domain context. Reflection means the agent evaluates its output against user intent before finalising. A system missing any one of these is not truly agentic.
What is the use-case-first principle?
The use-case-first principle states that every architectural decision — LLM model, architecture type, tools, prompts, guardrails — must flow from a clearly defined use case. Building an agent without knowing why is described as the most common and most costly mistake. Before touching any technology, you answer two questions: what friction am I removing, and what non-value-added activity am I releasing humans from?
What does 'friction' mean in the context of AI agents?
Friction refers to the number of steps a human must take to complete a process. Agents are built to reduce or eliminate friction by collapsing multi-step workflows into minimal interactions. When defining your use case, you identify which multi-step human process the agent will collapse — this is one of two core justifications for building the agent at all.
What are non-value-added activities and why do they justify an agent?
Non-value-added activities are tasks the organisation must perform but for which customers are not willing to pay — like routine status update emails or debugging known error patterns. They justify an agent because deploying automation here releases human capacity for work customers do value. Identifying these activities is the second core question in Step 1 of the framework.
// How To
How do I set up short-term memory in my agent?
Set up short-term memory by configuring your framework's message history handler to store the full chat conversation history for the current session. This lets the agent maintain context across multiple turns of dialogue. In LangGraph, Crew AI, or similar libraries, this is a built-in message history component — you enable and scope it to the active session.
How do I decide between RAG and a direct database connection for long-term memory?
Decide based on data structure. If your external data is unstructured — documents, code repositories, PDFs — use RAG Architecture to chunk, embed, and retrieve it. If your data lives in a structured database with clear schemas, connect directly without RAG. Long-term memory defines how domain-aware your agent will be, so match the retrieval method to your data type.
How do I engineer the system prompt for reflection behaviour?
Encode reflection in the system prompt by instructing the agent to evaluate its output against the user's original intent before finalising a response. Use a Reflection Prompt strategy, and combine it with the right prompting approach for your task — Zero-Shot, Few-Shot, Chain-of-Thought, or Tree-of-Thought. The system prompt governs the agent's identity, role, and behavioural constraints.
How do I test whether my agent is production-ready?
Run the agent against the use case criteria you defined in Step 1, ideally on a set of historical or sandboxed cases. Measure accuracy, escalation rate, and loop count. If it fails to satisfy the use case, return to the relevant layer — architecture, tools, prompts, or guardrails — and adjust. If it passes, deploy. Do not skip evaluation.
// Troubleshooting
My agent stops at the first error instead of retrying — how do I fix this?
This means you have skipped Loop Engineering. Without deliberately designed iterative feedback cycles, the agent cannot achieve autonomy and will halt at the first failure. Fix it by implementing loops that let the agent detect the error, re-run the autonomy-tool-memory cycle, and persist toward the goal — for example, re-running tests after each fix attempt before escalating to a human.
My agent gives generic, unusable outputs — what's wrong?
Your agent likely lacks long-term memory. Without feeding it organisational context via RAG Architecture or a direct database connection, it has no domain awareness and defaults to generic responses. Treating long-term memory as optional is a listed pitfall. Populate long-term memory with your domain data, code repos, or business rules so the agent operates with the context it needs.
My agent behaves unpredictably in production — what did I miss?
You likely omitted one or more guardrail layers. Guardrails apply at four distinct control points — Input, LLM, Tool, and Output — and leaving any one unset creates unpredictable behaviour. Audit all four: define which inputs the agent responds to, constrain model tone and escalation, limit what actions tools may take, and enforce output format and safety. Guardrails are not optional in production.
My multi-agent system is hard to debug — should I have used a single agent?
Probably yes. Choosing multi-agent architecture when a single agent will do is a common pitfall — it increases failure surface and debugging difficulty. The rule is to always test single-agent viability first. If a single ReAct or Plan-and-Execute agent can solve the use case, refactor down. Only keep multi-agent complexity when the task genuinely requires specialised delegating sub-agents.
// Comparisons
How does this framework compare to a generic 'connect an LLM to some tools' approach?
A generic approach often bolts tools onto an LLM without a use case, memory strategy, or guardrails — producing a fragile system that stops at errors. This framework enforces a disciplined 7-layer sequence starting from the use case, mandates all four agentic traits, and applies four-layer guardrails. The result is a production-ready agent with justified model choices, domain awareness, and reliable self-correction.
How does ReAct architecture compare to Plan and Execute?
ReAct alternates between reasoning and acting in a loop — the agent thinks, takes a tool action, observes, and repeats — which suits dynamic tasks where the next step depends on the last result. Plan and Execute produces a complete plan upfront, then executes each step sequentially, which suits predictable, well-scoped tasks. Both are single-agent patterns; choose based on how much the workflow branches at runtime.
How does Hierarchical multi-agent compare to Swarm agents?
Hierarchical agents use an orchestrator that delegates subtasks to specialised sub-agents, giving centralised control and clear responsibility per agent. Swarm agents operate in parallel with no central orchestrator, coordinating dynamically to solve a problem. Hierarchical suits tasks with clear delegation like diagnosis then fix; swarm suits problems benefiting from parallel exploration without a controlling authority.
When should I choose a mid-tier LLM over a flagship model?
Choose a mid-tier model like Sonnet or a mini variant when your use case involves high volume with moderate reasoning depth — the token billing savings compound at scale and improve ROI. Reserve flagship models like Opus or GPT-5 for complex reasoning tasks such as diagnosing root causes in a codebase. Never default to the largest model without cost justification tied to the use case.
// Advanced
Which framework should I build with — LangGraph, Crew AI, Microsoft Agent Framework, or LlamaIndex?
All four are valid Python libraries for building agentic architectures; choose based on your architecture pattern and ecosystem fit. LangGraph excels at graph-based and hierarchical flows with explicit state control. Crew AI suits role-based multi-agent collaboration. Microsoft Agent Framework (formerly AutoGen) supports conversational multi-agent setups. LlamaIndex is strong for retrieval-heavy, RAG-centric agents. Match the tool to your architecture choice from Step 3.
How do I apply guardrails to tools that write to production systems?
Apply Tool Guardrails that define exactly what actions each tool may take. For code-writing agents, limit writes to non-main branches only, as shown in the CI pipeline example. Combine this with Output Guardrails requiring a structured report on every action. Tool guardrails are one of four distinct control points and are essential for any agent that can mutate real-world state.
Can I use Chain-of-Thought and Reflection prompts together?
Yes — they serve different purposes and combine well. Use a Chain-of-Thought Prompt to guide the agent through step-by-step reasoning during diagnosis or planning, then use a Reflection Prompt to make the agent evaluate whether its output resolved the problem before finalising or escalating. The CI pipeline example uses exactly this pairing: CoT for diagnosis, reflection for re-running tests after each fix.
How do I justify the ROI of building an agent to stakeholders?
Justify ROI by grounding the build in the use-case-first principle: quantify the friction removed (steps collapsed in a multi-step process) and the non-value-added activity eliminated (work customers won't pay for). Then tie your LLM model selection to token billing at projected scale and the projected value of solving the use case. This makes cost and benefit explicit before any code is written.
How does Loop Engineering apply to my build process, not just runtime?
The Loop Engineering principle applies to your build process as well as the agent's runtime behaviour. In Step 8, you run the agent against use case criteria, and if it fails, you loop back to the relevant layer — architecture, tools, prompts, or guardrails — adjust, and re-test. Just as the agent self-corrects, your build should iterate until the output meets the required standard.
What is the minimum viable set of tools for an agent?
There is no fixed minimum — you map each tool to a specific action the use case requires, so the minimum is exactly the tools needed to perform the required actions and achieve autonomy. Common ones include web search for real-time data, a Python execution environment for running and fixing code, Gmail for email actions, and SQL databases for structured records. Add nothing that isn't tied to a use case action.