How to Integrate LLMs with SAP Without Breaking the Core
For Enterprise architects integrating LLMs with SAP/ERP · Based on Cloud Guru LLM Engineering Production Framework
// TL;DR
If you are integrating LLMs with SAP or another ERP, the Cloud Guru LLM Engineering Production Framework gives you an upgrade-safe pattern: build the agent side-by-side on BTP and reach into SAP through stable OData and RAP APIs, never hacking the frozen core. Ground every model answer in real SAP data, because a plausible-looking wrong number is a financial statement problem, not a bug. The LLM drafts — extraction, classification, natural language to query — and deterministic SAP logic commits. Gate every irreversible action on human approval and log every model call with full prompt and completion for audit. Use it when automating invoice intake, master-data workflows, or S/4HANA migration.
Why can't I just embed LLM calls inside my SAP system?
Because of the Clean Core principle: you do not hack the standard system or bury AI calls inside frozen code. If you do, the next SAP upgrade breaks your integration. Instead, expose and consume stable released APIs so the system can be upgraded underneath you without touching your AI logic.
The recommended architecture is the side-by-side pattern: build the AI agent as a separate application on BTP that reaches into SAP through OData and RAP. This keeps the ERP core clean and upgrade-safe while giving your agent everything it needs.
How do I stop the LLM from producing a plausible but wrong number?
Ground every model answer in real SAP data, and keep a hard division of labour: the LLM drafts, deterministic SAP logic commits. The model handles extraction, classification, and natural-language-to-query. SAP posts documents and moves money. A plausible-looking wrong number is not a cosmetic bug — it is a financial statement problem.
For an invoice-intake workflow, the agent receives the scanned PDF, calls the generative AI hub with a structured extraction prompt, and receives typed JSON — supplier, amount, PO number. Then apply the Constrain → Validate → Retry loop: validate the JSON against a schema, check the supplier against master data, and confirm the PO exists. On failure, send the error back to the model and retry.
Where do humans stay in the loop?
Gate every irreversible action on human approval. In invoice intake, the extracted, validated data is presented to a person before any deterministic ERP action posts the document. This is non-negotiable for anything that touches money or master data. And never autodeduplicate master data — flag candidate duplicates and present them to a human, because an automated merge on faulty logic corrupts the system of record.
How do I use RAG safely in an SAP context?
Use the HANA Cloud vector engine for RAG grounding so retrieval stays inside your governed data estate. Build the pipeline with the standard checklist: chunk on meaning, retrieve with hybrid search, rerank the short list, and evaluate retrieval and generation separately. Because enterprise queries often contain exact material codes and document numbers, hybrid search with BM25 is essential — pure semantic search will miss them.
How do I defend an enterprise agent against prompt injection?
Treat every byte the agent reads — an emailed invoice, a fetched document — as potentially hostile data, not commands. Keep a clear privilege boundary, require human confirmation for irreversible actions, and validate all outputs against schema and policy. Indirect prompt injection, where malicious text hides inside content the agent retrieves, is the dangerous variant, and there is no complete fix — only layered risk reduction.
How do I satisfy audit and compliance requirements?
Log every model call with the full prompt and completion, then apply redaction and retention policies rather than firehosing plain text. Trace requests as spans using OpenTelemetry GenAI semantic conventions so LLM calls sit in the same observability backend as the rest of your enterprise services. For a non-deterministic model in a financial process, the recorded input and output are your audit trail and your bug report at once.
Next step: Map one candidate workflow — invoice intake is ideal — into the draft-versus-commit split: list exactly which steps the LLM drafts and which steps deterministic SAP logic commits, then place a human approval gate before every commit. That boundary is your whole architecture in miniature.
// FREQUENTLY ASKED QUESTIONS
What is the safest architecture for LLMs with SAP?
The side-by-side pattern: build the AI agent as a separate application on BTP that reaches into SAP through stable OData and RAP APIs. This honours the Clean Core principle — you never modify the standard system or bury AI calls in frozen code — so SAP can upgrade the system underneath you without breaking your integration.
Can I let an LLM post documents directly in my ERP?
No. The LLM drafts — extraction, classification, natural language to query — and deterministic SAP logic commits, posting documents and moving money. Gate every irreversible action on human approval. A plausible-looking wrong number posted automatically is a financial statement problem, not a cosmetic bug.
Should I let an agent deduplicate master data automatically?
Never autodeduplicate master data. Flag candidate duplicates and present them to a human instead. An automated merge based on faulty model logic corrupts your system of record. Keep the LLM in a suggest-and-present role for anything touching master data, and require explicit human confirmation before any change.
How should I ground an SAP RAG system so it doesn't hallucinate numbers?
Use the HANA Cloud vector engine for grounding and build the pipeline with meaning-based chunking, hybrid search, and reranking. Hybrid search with BM25 is essential because enterprise queries carry exact material codes and document numbers that pure semantic search misses. Evaluate retrieval and generation separately, and validate every extracted value against real SAP master data before use.