Frequently Asked Questions About Srinivasan AI Engineer Stack Framework 2026
20 answers covering everything from basics to advanced usage.
// Basics
What are the six tool categories in the framework?
The six categories are: (1) Building and Orchestration—LangChain, LangGraph, OpenAI Agent SDK; (2) Connectivity and Enterprise Integration—MCP; (3) Models and Inference Runtimes—managed platforms, vLLM, Triton, BentoML; (4) Retrieval and Vector Databases—PGVector, Weaviate, Pinecone; (5) Evaluation Toolkits—LangSmith, RAGAS, TruLens, MLflow; (6) Observability—OpenTelemetry, OpenInference, LangSmith, Galileo, Phoenix/Arize.
What is MCP and why do beginners overlook it?
MCP (Model Context Protocol) is the standard governing how LLM applications connect to tools and data sources, defining authentication, permissions, and capability exposure. Beginners consistently skip the Connectivity and Enterprise Integration category, but understanding least privilege tool exposure and auditability is a major differentiator in enterprise environments. Learning MCP early sets you apart.
What's the difference between LangChain and LangGraph?
LangChain provides core abstractions—messages, prompt templates, runnables, output parsers, tool calling, structured outputs, and retrieval integrations. LangGraph is what you graduate to when you need durable, stateful agent workflows: graph execution, checkpointing, human-in-the-loop control points, and planner/executor/validator separation. Start with LangChain, move to LangGraph once basic agent patterns feel comfortable and durability becomes the bottleneck.
// How To
Why should I study the OpenAI Agent SDK even if I use a different model provider?
Study the OpenAI Agent SDK regardless of model provider because it's a masterclass in clean agent architecture: explicit tools, guardrails, handoffs between agents, and typed outputs. The architectural discipline transfers to any framework or provider. It teaches you how to design agents properly, which is more valuable long-term than provider-specific API details.
How do I audit my existing AI system for gaps?
Run a lifecycle audit: place every tool you use against the phases Build → Orchestrate → Retrieve → Run → Evaluate → Observe → Iterate. Gaps usually appear in evaluation (Category 5) and observability (Category 6). If your RAG hallucinates, revisit Category 4 for chunking strategy and hybrid search. If agent logic breaks unpredictably, you may need LangGraph for stateful durability.
How do I build my first RAG pipeline with this framework?
Start with PGVector on Postgres if you already use it—it's the pragmatic embedded choice. Learn index and distance metric choices, hybrid retrieval patterns, and chunking failure modes. Always address the context assembly problem: chunking strategy and hybrid search configuration are where most RAG systems fail. Add RAGAS from the start to measure faithfulness, relevance, and groundedness.
What's the minimum viable stack to demonstrate in AI engineering interviews?
A full-lifecycle portfolio without overwhelming self-hosted infrastructure: LangChain plus LangGraph for building and orchestration, MCP concepts for enterprise credibility, managed inference trade-offs, a simple RAG pipeline on PGVector, LangSmith evaluation framed as CI from day one, and OpenTelemetry basics for observability. This proves you understand how tools connect across the entire lifecycle.
// Troubleshooting
My RAG system keeps hallucinating. How do I fix it?
Hallucinations usually trace to the context assembly problem, not the retrieval mechanism. Introduce RAGAS to measure faithfulness, relevance, and groundedness. Use TruLens for component-level scoring to isolate whether failures are retrieval or generation failures. Add OpenTelemetry and OpenInference tracing for trace-level visibility. Then audit your chunking strategy and hybrid search configuration—the likely root cause.
My AI agent breaks unpredictably in production. What's wrong?
Unpredictable agent behavior often means basic LangChain patterns are your durability bottleneck. Move to LangGraph for durable, stateful workflows with checkpointing and human-in-the-loop control points. Separate planner, executor, and validator into distinct inspectable graph nodes. Add trace-level debugging via OpenInference so you can inspect prompts, retrieval steps, tool calls, and model invocations when something goes wrong.
I'm learning too many tools and none of them deeply. How do I focus?
Stop learning every tool that appears—that's the top pitfall. Anchor to your specific goal within the AI application lifecycle first. Learn one tool per phase deeply before adding another. Depth on a curated set beats shallow familiarity with everything. If you can't place a tool in the lifecycle, deprioritise it entirely.
// Comparisons
How does this framework compare to a generic AI bootcamp?
Generic bootcamps often teach a fixed curriculum regardless of your background or deployment context. This framework is personalised—it maps recommended tools against the lifecycle based on your role, goal, existing knowledge, and whether you're building for startups or enterprise. It also enforces evaluation and observability discipline that most bootcamps treat as optional afterthoughts.
How does managed inference compare to self-hosted inference?
Managed inference (calling APIs) is best for rapid development—you learn latency, cost, quality, and context-length trade-offs plus concurrency, streaming, batching, embeddings, and re-rankers. Self-hosted inference (vLLM, Triton, BentoML) is for optimising throughput and controlling costs—you learn KV caching constraints, quantization trade-offs, and tail latency. Serious AI engineers need both sides of inference.
Should I use Weaviate or Pinecone for my vector database?
Choose based on deployment context. Weaviate is open-source with built-in hybrid search—good when you want control or to self-host. Pinecone is managed and built for production at scale—good when you want to offload infrastructure. The framework recommends learning one embedded option (PGVector) plus one dedicated vector database, since schema design, ingestion, and index trade-offs are universal.
What's the difference between LangSmith, RAGAS, and TruLens for evaluation?
LangSmith builds datasets from real traces and runs offline vs online eval with rule-based, LLM-as-a-judge, or human review evaluators. RAGAS focuses on RAG quality metrics—faithfulness, relevance, groundedness—ideal for regression testing. TruLens scores individual components (retrieval, tool use, final answer) separately using a feedback function style. Use them together for full-coverage evaluation.
// Advanced
How do I implement enterprise observability discipline?
Start with OpenTelemetry as your vendor-neutral foundation: traces, metrics, logs, context propagation, latency percentiles, and error budgets. Add OpenInference for LLM-specific span attributes covering prompts, retrieval, tool calls, and model invocations. Then adopt one or two LLM-native platforms like LangSmith, Galileo, or Phoenix/Arize. Master trace-level debugging, comparing releases on latency/cost/quality, and linking eval regressions to concrete traces.
What is the planner/executor/validator separation pattern?
It's a LangGraph architectural pattern where an agent's planning logic, execution logic, and validation logic are kept as distinct, inspectable graph nodes rather than collapsed into a single agent loop. This separation makes agents debuggable and durable, supports checkpointing and human-in-the-loop control, and prevents the opaque failures common in monolithic agent designs.
When should I learn Nvidia Triton versus vLLM?
Learn vLLM for serving open-source LLMs—it's the go-to engine where you master throughput vs tail latency, KV caching constraints, memory behavior, and quantization trade-offs. Learn Triton Inference Server for non-LLM models like embeddings, re-rankers, and computer vision, where you master dynamic batching, concurrency tuning, model repositories, and ensembles. BentoML complements both for packaging and deployment discipline.
How do I use MLflow for GenAI evaluation?
Use MLflow to treat prompt runs and evaluations like tracked ML artifacts. Version your prompts, compare runs quantitatively, and store evaluation artifacts per release. This brings ML experiment-tracking rigor to GenAI, letting you link a quality regression back to a specific prompt or release—core to treating evaluation as CI for AI behavior.
What deployment context factors change which tools I should learn?
Personal projects favor managed APIs and embedded stores like PGVector for speed. Startups balance managed inference with cost-aware choices. Enterprise environments demand MCP for least privilege and auditability, self-hosted inference for cost control, dedicated vector databases with multi-tenant patterns, and rigorous observability with OpenTelemetry plus enterprise platforms like Galileo. Always match tool priority to context.
Is it worth learning fine-tuning in 2026?
Learn fine-tuning selectively—if your work requires it, the key skill is versioning discipline, treating fine-tuned models as tracked artifacts. For most rapid development, managed platform trade-offs (latency, cost, quality, context length) and retrieval via RAG solve the problem without fine-tuning. Prioritise fine-tuning only after you've covered building, retrieval, evaluation, and observability.