AI Engineering Interview Guide — Part 4: Production AI Architecture
How to turn an LLM/RAG/Agent POC into a secure, scalable, observable, cost-efficient enterprise AI platform
In the previous three articles, we built the conceptual foundation:
Part 1 — LLM Fundamentals
We learned about tokens, embeddings, vectors, context windows, model parameters, and fine-tuning.
Part 2 — RAG Deep Dive
We explored ingestion, chunking, embeddings, retrieval, reranking, metadata, evaluation, and RAG security.
Part 3 — Agentic AI
We looked at agents, tools, planning, memory, orchestration, autonomy, and agent security.
Now comes the question that separates a prototype from a real production system:
"How would you architect this for an enterprise?"
A POC can often be built with:
Python
+
LLM API
+
Simple vector store
+
A notebook
Production is very different.
Now you need to think about:
Security
Scalability
Availability
Latency
Cost
Data governance
Multi-tenancy
Observability
Failure handling
Versioning
Disaster recovery
Compliance
And the architecture becomes something more like:
USERS
│
▼
┌─────────────┐
│ API / Web │
│ Application │
└──────┬──────┘
│
▼
Identity / IAM
│
▼
AI Gateway Layer
│
┌──────────────┼──────────────┐
▼ ▼ ▼
LLM RAG Agents
│ │ │
│ ┌────┴────┐ Tool Gateway
│ ▼ ▼ │
│ Search Vector DB │
│ │
└──────────────┬───────────────┘
▼
Enterprise Data / APIs
│
▼
Observability / Audit
This article focuses on architecture decisions, not framework tutorials.
1. The First Architecture Question: What Are We Actually Building?
One of the biggest mistakes in AI architecture is starting with:
"Which LLM should we use?"
That's too early.
Start with the workload.
For example:
Business Requirement
│
▼
What does the system actually do?
│
├── Generate text?
├── Answer questions?
├── Search enterprise data?
├── Make decisions?
├── Execute actions?
└── Perform multi-step tasks?
These map to different architectures.
Simple generation
User → LLM → Response
Enterprise knowledge assistant
User → RAG → LLM
Autonomous task assistant
User → Agent → Tools → Enterprise Systems
Complex enterprise AI platform
User
↓
AI Application
↓
RAG + Agents + APIs + Enterprise Data
The first architectural principle is therefore:
Design around the workload and risk, not around the model.
2. POC Architecture vs Production Architecture
Let's take a practical example.
Imagine an engineer creates a security assistant POC:
User
↓
Python script
↓
LLM API
↓
FAISS
↓
PDF files
It works.
The demo looks great.
Then someone says:
"Let's make this available to 20,000 employees."
The architecture changes immediately.
You now need:
USERS
│
▼
API Gateway
│
Authentication
│
▼
Application Layer
│
┌──────────┼──────────┐
▼ ▼ ▼
RAG Agent LLM
│ │
▼ ▼
Search Tools
│ │
▼ ▼
Data Enterprise APIs
And around it:
IAM
Secrets
Observability
Rate Limiting
Caching
Audit
Security Controls
CI/CD
A useful interview answer is:
"My POC architecture optimizes for learning and speed. Production architecture optimizes for security, reliability, scale, cost, operability, and governance."
3. A Reference Enterprise AI Architecture
Let's build a reusable reference architecture.
┌───────────────┐
│ Users │
└───────┬───────┘
│
▼
┌──────────────────┐
│ API / Web Layer │
└────────┬─────────┘
│
▼
┌────────────────────┐
│ Identity / Access │
└─────────┬──────────┘
│
▼
┌─────────────────────┐
│ AI Gateway / │
│ Application │
└─────────┬───────────┘
│
┌──────────────────┼──────────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌───────────┐
│ LLM │ │ RAG │ │ Agent │
└──────────┘ └────┬─────┘ └─────┬─────┘
│ │
┌──────┴──────┐ Tool Gateway
▼ ▼ │
Search Vector Store │
│
┌───────────────────┼────────────────┐
▼ ▼ ▼
Jira GitHub DB
Supporting services:
Secrets Management
Policy / Authorization
Cache
Observability
Audit Logging
DLP
Rate Limiting
Configuration
Evaluation
The exact technologies will change from company to company.
The architectural responsibilities should not.
4. Separate the AI Layer From the Application Layer
An enterprise AI system should not become:
Frontend
↓
LLM
↓
Everything
Instead, separate responsibilities.
Application
│
▼
AI Orchestration
│
┌──────────────┼──────────────┐
▼ ▼ ▼
LLM RAG Tools
Why?
Because the application must control:
authentication
authorization
business rules
validation
transactions
rate limits
logging
error handling
The LLM should not own these responsibilities.
A useful architectural principle is:
Keep probabilistic reasoning isolated from deterministic system controls.
5. Introduce an AI Gateway
If many applications directly call different model providers:
App A → Provider A
App B → Provider B
App C → Provider C
App D → Provider A
model management becomes difficult.
A common enterprise pattern is:
Applications
│
▼
AI Gateway
│
┌───┼───────────────┐
▼ ▼ ▼
LLM A LLM B Local Model
The gateway can centralize:
model routing
authentication
quotas
rate limiting
logging
policy enforcement
cost tracking
provider abstraction
model version management
This is especially useful when the organization uses multiple models.
6. Why Model Abstraction Matters
Imagine your application directly depends on a provider-specific API.
Later:
The provider changes the model.
or:
A cheaper model becomes available.
or:
The organization wants an on-prem model.
If every application is tightly coupled to one provider, migration becomes painful.
A model abstraction layer can provide:
Application
│
▼
Model Interface
│
┌───┼──────────────┐
▼ ▼ ▼
Model A Model B Local Model
This doesn't mean every model should behave identically.
It means the application should avoid unnecessary provider coupling.
7. Model Selection Is an Architecture Decision
Don't ask:
"Which model is the best?"
Ask:
"Which model is appropriate for this workload?"
Consider:
Quality
Latency
Cost
Context size
Reasoning capability
Multilingual support
Privacy requirements
Data residency
Availability
Rate limits
Operational maturity
For example:
Simple classification
↓
Smaller model
Complex reasoning
↓
Larger model
Sensitive workload
↓
Approved private deployment
This naturally leads to model routing.
8. Model Routing
Suppose an enterprise platform receives 1 million AI requests every day.
Not all requests are equally complex.
A router can classify the request:
Request
│
▼
Task Router
│
┌──────────┼──────────┐
▼ ▼ ▼
Simple Medium Complex
│ │ │
▼ ▼ ▼
Model A Model B Model C
This can reduce cost while maintaining quality.
But the router itself should be measured.
Don't assume:
"Smaller model = always good enough."
Use evaluation data to validate the trade-off.
9. Data Architecture
AI systems often touch many different kinds of data.
For example:
Documents
Structured records
Embeddings
User conversations
Agent state
Application data
Logs
Telemetry
Model outputs
Evaluation data
Trying to put everything into one database is usually a mistake.
A better design might be:
Data Layer
│
┌───────────────┼────────────────┐
▼ ▼ ▼
Object Storage Relational DB Vector Store
│ │ │
Documents Business Data Embeddings
And possibly:
Cache
Search Index
Graph Store
Event Store
Choose the storage technology based on the workload.
10. Object Storage vs Database vs Vector Store
A simple rule of thumb:
Object storage
Good for:
PDFs
documents
large files
raw datasets
Relational database
Good for:
transactions
users
metadata
configuration
business entities
Vector store
Good for:
embeddings
semantic retrieval
Search engine
Good for:
keyword retrieval
filtering
analytics-style search
Cache
Good for:
frequently accessed data
expensive computations
low-latency responses
The architecture should reflect the data's access pattern.
11. RAG in Production Is a Data Pipeline
A production RAG architecture isn't simply:
PDF → Vector DB
It is more like:
Source Systems
│
▼
Ingestion
│
▼
Parse / Normalize
│
▼
Classify
│
▼
Chunk
│
▼
Metadata
│
▼
Embed
│
▼
Index / Publish
│
▼
Search Layer
This should behave like a proper production data pipeline.
You need:
retries
dead-letter handling
idempotency
monitoring
versioning
change detection
delete propagation
12. Event-Driven Document Ingestion
Instead of periodically scanning everything:
Every night
↓
Scan all documents
you can use events:
Document Updated
│
▼
Event
│
▼
Ingestion Worker
│
▼
Parse → Chunk → Embed → Index
This makes updates more responsive and can reduce unnecessary work.
The same pattern applies to deletion:
Document Deleted
↓
Deletion Event
↓
Remove corresponding chunks
↓
Invalidate cache/index
13. Idempotency Is Important in AI Pipelines
Suppose the document ingestion worker fails halfway through.
A retry happens.
Without idempotency you might create:
Document A
├── Chunk 1
├── Chunk 2
├── Chunk 3
├── Chunk 1 ← duplicate
├── Chunk 2 ← duplicate
└── Chunk 3 ← duplicate
Use stable identifiers.
For example:
document_id
document_version
chunk_id
so the pipeline can safely retry.
AI systems inherit many of the same distributed-systems problems as conventional applications.
14. Multi-Tenancy Must Be Designed, Not Added Later
Suppose your platform supports:
Tenant A
Tenant B
Tenant C
You need a clear answer to:
Where does tenant isolation happen?
Potential boundaries include:
Tenant
↓
Application
↓
Database
↓
Vector Index
↓
Cache
↓
Logs
One common logical model is:
tenant_id = A
being propagated through the entire request.
User
↓
Tenant Context
↓
RAG Retrieval
↓
Tools
↓
Database
This becomes extremely important because AI systems often combine information from multiple sources.
15. Cache Design Can Create Security Problems
Caching is useful.
For example:
User Query
↓
Cache
↓
Cached response
But imagine:
Tenant A asks:
"What is our architecture?"
and the response is cached using only:
query = "What is our architecture?"
Then Tenant B asks the same question.
Oops.
The cached answer could be returned to the wrong tenant.
The cache key may need to include relevant security context:
tenant_id
user_scope
authorization context
query
This is a good example of a broader principle:
Every optimization layer must preserve the security model.
16. Authorization Must Flow Through the Architecture
Suppose a user asks an agent:
"Show me all security tickets."
The request may travel through:
User
↓
Application
↓
Agent
↓
Jira Tool
↓
Jira
The authorization context must not disappear along the way.
Conceptually:
User Identity
│
▼
Authorization Context
│
▼
Agent
│
▼
Tool Gateway
│
▼
Enterprise Resource
The agent may be autonomous, but it still operates within a user's security context.
17. Don't Let Business Rules Live in the Prompt
Imagine a business rule:
"Only security managers can approve production exceptions."
Don't implement this as:
System Prompt:
If user is not a security manager,
do not approve.
That is useful as guidance but weak as enforcement.
Instead:
Agent
↓
Approval Request
↓
Authorization Service
↓
Policy
↓
ALLOW / DENY
The deterministic policy system owns the business rule.
18. Human Approval Architecture
For sensitive actions, introduce an approval boundary.
Example:
Agent
│
▼
Proposed Action
│
▼
Risk Engine
│
┌──────┴──────┐
▼ ▼
Low Risk High Risk
│ │
▼ ▼
Execute Human Approval
│
┌────┴────┐
▼ ▼
Approve Reject
│
▼
Execute
This is especially useful for:
production changes
financial operations
destructive actions
external communication
privileged access changes
19. Design for Failure
Production AI systems will fail.
A model might be unavailable.
A vector store might be unavailable.
A tool may time out.
A provider may throttle requests.
An architecture should explicitly define what happens.
Request
│
▼
Agent
│
▼
Tool
│
Failure
│
┌─────────┼─────────┐
▼ ▼ ▼
Retry Fallback Escalate
But retries require care.
20. Retry Isn't Always Safe
Consider:
create_payment()
The request times out.
You don't know whether the payment succeeded.
Retrying may create a duplicate transaction.
Therefore:
Retries need idempotency semantics, especially for write operations.
A useful architecture includes:
Request ID
Idempotency Key
Timeout
Retry Policy
Circuit Breaker
This is a distributed-systems lesson that applies directly to agentic AI.
21. Timeouts and Cancellation
Imagine an agent starts:
Tool A → 10 sec
Tool B → 15 sec
Tool C → 30 sec
Tool D → hangs
Without controls, the user may wait indefinitely.
Set:
Request timeout
Tool timeout
LLM timeout
Maximum agent duration
And make cancellation possible.
For an agent system:
Cancel Request
↓
Orchestrator
↓
Stop pending tool calls
↓
Terminate execution
22. Agent Budgets
Agents can consume unpredictable resources.
Set explicit limits:
Maximum steps
Maximum tool calls
Maximum tokens
Maximum execution time
Maximum estimated cost
For example:
MAX_STEPS = 15
MAX_TOOL_CALLS = 25
MAX_RUNTIME = 2 minutes
The numbers are workload-specific.
The architectural principle is what matters:
Every autonomous loop should have a budget.
23. Observability Is Not Optional
In a traditional API, you might log:
Request
Response
Latency
Status
For an agentic system, that's not enough.
You may need to understand:
User request
↓
LLM call
↓
Tool selection
↓
Tool parameters
↓
Tool result
↓
Next LLM call
↓
Another tool
↓
Final answer
This is why distributed tracing becomes especially valuable.
24. Trace the Entire AI Request
A useful trace might look like:
Request ID: 12345
├── Authentication
├── Query processing
├── Embedding generation
├── Retrieval
│ ├── Vector search
│ └── Reranking
├── LLM call #1
│ └── Tool: get_jira_ticket
├── Jira API
├── LLM call #2
│ └── Tool: get_github_pr
├── GitHub API
├── LLM call #3
└── Final response
Now an engineer can answer:
Why did this request take 18 seconds?
or:
Why did the agent call Jira four times?
25. What Should You Measure?
A mature AI platform has multiple dimensions of observability.
Infrastructure
CPU
Memory
Network
Storage
AI
Token usage
Latency
Model errors
Context size
RAG
Retrieval latency
Recall@K
No-result rate
Reranking latency
Agent
Steps
Tool calls
Tool failures
Task completion
Business
Success rate
User satisfaction
Escalation rate
Security
Policy violations
Blocked actions
Sensitive-data detections
Unauthorized access attempts
26. Cost Architecture
AI systems can become expensive surprisingly quickly.
Consider:
1 million requests/day
×
5,000 input tokens
×
large model
That can become a significant operational cost.
Cost should therefore be designed into the architecture.
Cost Control
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Caching Model Routing Retrieval
│ │ │
▼ ▼ ▼
Fewer calls Cheaper models Less context
27. Token Economics
Consider an agent workflow:
Step 1 → 3,000 tokens
Step 2 → 4,000 tokens
Step 3 → 5,000 tokens
Step 4 → 4,000 tokens
Step 5 → 6,000 tokens
The system may process tens of thousands of tokens for one task.
Cost optimization therefore isn't only:
"Use a cheaper model."
You should also ask:
Why are we sending so much context?
Potential improvements:
better retrieval
shorter prompts
context compression
caching
fewer agent loops
smaller models for simple steps
structured outputs
28. Caching Strategy
Different things can be cached.
Embeddings
Avoid generating the same embedding repeatedly.
Retrieval results
Useful when queries repeat and the underlying data hasn't changed.
LLM responses
Useful for deterministic or low-volatility workloads.
Tool results
Potentially useful for safe, short-lived data.
But every cache needs to respect:
Freshness
Authorization
Tenant isolation
Data sensitivity
29. Semantic Caching
Traditional caching depends on exact keys.
"How do I reset my password?"
Semantic caching can recognize:
"How can I change my forgotten password?"
as potentially similar.
Conceptually:
New Query
↓
Semantic Similarity
↓
Existing Cached Query
↓
Reuse result?
This can reduce model calls.
But semantic caching is risky for sensitive workloads if the cache scope and authorization model aren't designed carefully.
30. Latency Architecture
An AI request might involve:
Authentication
+
Embedding
+
Retrieval
+
Reranking
+
LLM
+
Tool Calls
Latency can accumulate.
A simple mental model is:
Total Latency
≈
Network
+
Retrieval
+
Reranking
+
LLM
+
Tool Calls
This is why architecture needs to distinguish between:
sequential calls
parallel calls
asynchronous work
streaming
31. Parallel Tool Execution
Suppose an agent needs:
Jira ticket
GitHub PR
Security policy
If these are independent:
Agent
├── Jira
├── GitHub
└── RAG
they may be executed in parallel rather than:
Jira → GitHub → RAG
Parallel execution can reduce latency.
But only when:
operations are independent
resource limits are respected
authorization is enforced consistently
32. Streaming Responses
Users don't necessarily need to wait for the entire response.
Instead:
LLM
↓
Token stream
↓
User sees partial response
This can improve perceived latency.
For agentic systems, however, don't confuse:
streaming intermediate text
with:
streaming dangerous actions.
Actions should remain governed and auditable.
33. Asynchronous Architecture
Some AI tasks don't need synchronous responses.
For example:
"Analyze 10,000 security findings and generate a report."
Instead of keeping an HTTP request open:
User
↓
Submit Job
↓
Queue
↓
Worker
↓
Agent / LLM
↓
Store Result
↓
Notify User
This is often much more resilient.
34. Queue-Based Architecture
For long-running AI jobs:
API
│
▼
Queue
│
┌─────────┼─────────┐
▼ ▼ ▼
Worker A Worker B Worker C
│ │ │
└─────────┼─────────┘
▼
AI Services
Benefits include:
load smoothing
retry management
backpressure
asynchronous processing
worker scaling
35. Backpressure
Imagine:
10 requests/sec
suddenly becomes:
10,000 requests/sec
Your LLM provider may not handle that load.
A queue allows the application to absorb bursts.
Users
↓
API
↓
Queue
↓
Controlled workers
↓
LLM
This protects downstream systems.
36. Rate Limiting
Rate limiting should exist at multiple levels.
User
↓
Tenant
↓
Application
↓
Model
↓
Tool
For example:
User:
100 requests/hour
Tenant:
10,000 requests/hour
Tool:
500 calls/minute
These are examples only.
The correct values depend on the business requirements.
37. Disaster Recovery
What happens if your AI platform loses:
Vector DB
Database
Object storage
Configuration
Agent state
You need recovery plans.
For critical systems:
Primary
│
├── Backups
├── Replication
└── Recovery procedures
For RAG specifically, remember:
The source documents are the authoritative data. Embeddings and indexes can often be rebuilt from them.
This can simplify disaster recovery planning.
38. Rebuild vs Backup
This is an interesting architecture decision.
Suppose the vector index is destroyed.
If you have:
Source documents
+
Chunking configuration
+
Embedding model/version
+
Metadata
you may be able to reconstruct it.
Therefore, the question becomes:
Is the vector index a primary data store or a derived index?
In many architectures:
Authoritative Source
↓
Derived
↓
Embedding / Index
This distinction can simplify recovery.
39. Version Everything
AI systems change frequently.
You should consider versioning:
Model
Prompt
Embedding model
Chunking strategy
Reranker
Knowledge index
Tool schemas
Agent policy
Evaluation dataset
For example:
Model: model-v4
Prompt: prompt-v17
Embedding: embed-v3
Index: index-2026-09
Policy: policy-v8
Now when quality changes, you can investigate what changed.
40. Prompt Versioning Is Engineering
Prompts should not live only in someone's laptop.
Think of prompts as software artifacts.
Prompt
↓
Git
↓
Review
↓
Test
↓
Deploy
↓
Monitor
A production prompt change should be treated similarly to code changes.
41. Evaluation Before Deployment
Suppose you want to replace:
Model A
with:
Model B
Don't simply deploy it.
Run an evaluation set:
Evaluation Dataset
│
┌──────┴──────┐
▼ ▼
Model A Model B
│ │
└──────┬──────┘
▼
Compare
Compare:
correctness
groundedness
latency
cost
safety
tool behavior
This is especially important because model changes can affect security behavior too.
42. Shadow Testing
A useful production technique is to send the same request to a new model without using its result operationally.
User Request
│
├── Production Model → User
│
└── Candidate Model → Evaluation Only
Then compare results.
This reduces migration risk.
43. Canary Releases
Instead of moving all users immediately:
100% Model A
you can introduce:
95% → Model A
5% → Model B
Monitor:
Error rate
Latency
Cost
Quality
Security events
Then gradually increase the percentage.
The same principle works for:
prompts
retrieval pipelines
agent policies
tool versions
44. AI CI/CD
A mature AI platform should treat AI changes as deployable artifacts.
A simplified pipeline:
Code / Prompt / Model Change
│
▼
Unit Tests
│
▼
AI Evaluation
│
▼
Security Test Suite
│
▼
Integration Tests
│
▼
Canary Deploy
│
▼
Production
This is essentially CI/CD for AI systems.
45. Traditional DevSecOps Still Matters
AI does not eliminate traditional security engineering.
The application still has:
APIs
Containers
Kubernetes
Databases
Identity
Secrets
Dependencies
Networks
Cloud infrastructure
Therefore, you still need:
SAST
DAST
SCA
Secrets scanning
Container scanning
API security
IAM review
Infrastructure security
And add AI-specific testing:
Prompt injection
RAG poisoning
Agent tool abuse
Data leakage
Excessive agency
The result is:
Traditional AppSec
+
AI Security
46. Secrets Management
Never put:
LLM_API_KEY=...
inside:
source code
prompt
agent memory
RAG documents
logs
Use dedicated secret-management mechanisms.
Conceptually:
Application
│
▼
Secret Manager
│
▼
Temporary Credential
│
▼
LLM / Tool
Rotate secrets and limit their scope.
47. Network Architecture
An enterprise AI platform may contain:
Internet-facing API
Internal RAG services
Vector database
Model endpoints
Tool servers
Databases
These shouldn't all live in the same unrestricted network.
A simplified model:
Internet
│
▼
API Gateway
│
Public Zone
│
▼
AI Application
│
┌───────────┼───────────┐
▼ ▼ ▼
RAG Tool Layer LLM
│ │
▼ ▼
Data Layer Enterprise APIs
Apply network segmentation and restricted egress where appropriate.
48. Observability Data Is Sensitive
There is a trap here.
An engineer might say:
"We'll log every prompt and every response for debugging."
Sounds useful.
But those logs may contain:
PII
Secrets
Customer data
Source code
Confidential documents
So AI observability requires:
Redaction
Classification
Access Control
Retention
Encryption
The debugging system itself becomes part of the security architecture.
49. Data Retention
Ask:
How long should the system retain?
Potentially:
Prompt
Response
Conversation
Agent trace
Retrieved documents
Tool results
Evaluation results
You may not want all of them retained forever.
A production design should define:
Purpose
Retention Period
Access
Deletion
Legal / Compliance Requirements
50. Model and Data Residency
For enterprise workloads, ask:
Where does the data physically or logically go?
A request might travel:
India
↓
Application
↓
US Model Provider
↓
US Logs
That may or may not be acceptable depending on organizational requirements.
Model architecture should therefore consider:
geographic residency
provider data handling
contractual controls
encryption
private endpoints
local models
This is an architecture requirement, not merely a compliance checkbox.
51. Open-Source Models vs Hosted Models
Both have trade-offs.
Hosted model
Advantages:
Fast adoption
Managed infrastructure
Less operational burden
Trade-offs may include:
Provider dependency
Data-handling considerations
External availability
Cost
Self-hosted model
Advantages:
More control
Potentially stronger isolation
Customization
Trade-offs:
GPU infrastructure
Operations
Scaling
Patching
Model lifecycle
Cost
Don't ask:
"Which is better?"
Ask:
"Which operational model fits our security, cost, performance, and data requirements?"
52. Kubernetes Architecture
For organizations already operating Kubernetes, an AI platform may look like:
Kubernetes Cluster
│
┌────────────────┼────────────────┐
▼ ▼ ▼
AI Gateway RAG Service Agent Service
│ │ │
▼ ▼ ▼
Model GW Vector DB Tool Gateway
But don't assume everything must run inside the cluster.
For example:
Kubernetes
│
├── Application
├── RAG
└── Agent
│
▼
Managed Model Service
Architecture should follow operational requirements.
53. Scaling Strategy
Different components scale differently.
Component Scaling Pattern
API Horizontal
RAG workers Horizontal
Embeddings Batch / worker scaling
LLM Provider-dependent
Vector DB Index/compute scaling
Agents Worker scaling
Tools Per-service scaling
Don't scale everything identically.
The bottleneck may be different at each layer.
54. A Common Architecture Mistake: Scaling the Wrong Layer
Suppose:
API servers = 100
but:
LLM quota = exhausted
Adding more API instances doesn't solve the actual bottleneck.
Likewise:
10 agent workers
may be useless if:
Tool API = 50 requests/minute
Production architecture requires capacity planning across the entire dependency chain.
55. Reliability Budget
A useful architectural mindset is to think in terms of:
Availability
Latency
Cost
Quality
Security
These often conflict.
For example:
More retrieval
↓
Better recall
↓
More tokens
↓
Higher cost
↓
Higher latency
Or:
Bigger model
↓
Potentially better quality
↓
Higher cost
↓
Higher latency
Architecture is therefore about managing trade-offs.
56. AI Quality Is a Production Concern
Traditional monitoring might say:
HTTP 200
But the AI answer could still be terrible.
Therefore, production monitoring should include quality signals.
For RAG:
Groundedness
Retrieval quality
No-answer rate
For agents:
Task completion
Tool accuracy
Unnecessary steps
Policy violations
This is one of the biggest differences between AI systems and traditional APIs.
57. AI SRE Mindset
A production AI team should ask:
"Is the system healthy?"
not simply:
"Is the HTTP endpoint up?"
Health should include:
Infrastructure
+
Model availability
+
Retrieval quality
+
Agent behavior
+
Latency
+
Cost
+
Security
You can have:
99.99% API uptime
and still have a completely unusable AI system because retrieval quality collapsed.
58. A Production AI Incident Example
Imagine users report:
"The assistant suddenly started giving incorrect security answers."
Infrastructure is healthy.
API is healthy.
Model endpoint is healthy.
What changed?
Yesterday:
Embedding Model v2
Today:
Embedding Model v3
The new embedding model changed the retrieval distribution.
Now:
Recall@5
92% → 61%
This is why AI observability must connect:
Model Version
+
Embedding Version
+
Index Version
+
Prompt Version
+
Quality Metrics
Without those links, debugging becomes guesswork.
59. Architecture for Safe Model Upgrades
A robust process is:
New Model
│
▼
Offline Evaluation
│
▼
Security Evaluation
│
▼
Shadow Traffic
│
▼
Canary
│
▼
Production
This is much safer than:
New model → deploy everywhere
60. A Full Production AI Platform
Let's bring everything together.
USERS
│
▼
┌────────────────┐
│ API / Web App │
└───────┬────────┘
│
▼
┌──────────────────┐
│ Identity / IAM │
└────────┬─────────┘
│
▼
┌─────────────────────┐
│ AI Gateway │
│ │
│ Routing │
│ Rate Limits │
│ Policy │
│ Cost Controls │
└──────────┬──────────┘
│
┌────────────────────┼─────────────────────┐
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌──────────┐
│ LLM │ │ RAG │ │ Agent │
└────────┘ └───┬────┘ └────┬─────┘
│ │
┌─────┴─────┐ Tool Gateway
▼ ▼ │
Search Vector DB │
│
┌─────────────┼─────────────┐
▼ ▼ ▼
Jira GitHub DB
┌────────────────────────────────────────┐
│ Data / Platform Layer │
│ │
│ Object Storage | SQL | Cache | Events │
└────────────────────────────────────────┘
┌────────────────────────────────────────┐
│ Governance / Security │
│ │
│ IAM | DLP | Secrets | Audit | Policy │
│ Network | Monitoring | Evaluation │
└────────────────────────────────────────┘
This is not one fixed architecture.
It's a reference model.
61. Architecture Decision: Centralized vs Distributed AI Platform
An enterprise may have two approaches.
Centralized
All AI applications
│
▼
Central AI Platform
Advantages:
common controls
model governance
centralized observability
consistent security
Distributed
Application A → AI stack A
Application B → AI stack B
Application C → AI stack C
Advantages:
application-specific optimization
team autonomy
Trade-off:
duplicated infrastructure
inconsistent controls
fragmented governance
A common enterprise compromise is:
Central platform capabilities
+
Application-specific logic
62. What Should Be Centralized?
Good candidates for centralization often include:
Model Gateway
Identity
Secrets
Policy
Audit
Observability
Evaluation infrastructure
Common retrieval infrastructure
Application-specific components may include:
Business workflow
Domain prompts
Domain tools
Domain knowledge
This gives teams flexibility without allowing every team to reinvent security and platform controls.
63. What Should Be Stateless?
Whenever practical, keep request-serving components stateless:
API
Agent workers
RAG service
Persist state externally:
Database
Cache
Queue
Object storage
This simplifies horizontal scaling.
For example:
Load Balancer
│
┌─────────┼─────────┐
▼ ▼ ▼
Worker A Worker B Worker C
│ │ │
└─────────┼─────────┘
▼
Shared State
64. What Should Be Durable?
Not everything should live only in memory.
Durable information might include:
User / tenant configuration
Agent state
Task state
Audit records
Source documents
Evaluation results
Business transactions
This is especially important when an agent executes a task over several minutes or hours.
65. Long-Running Agent Architecture
For complex tasks:
User
↓
Create Job
↓
Queue
↓
Agent Worker
↓
Checkpoint State
↓
Tool Calls
↓
Checkpoint State
↓
More Tool Calls
↓
Complete
If Worker A crashes:
Worker B
↓
Resume from checkpoint
This is much more robust than storing everything in a process's memory.
66. Security Architecture: The Three Most Important Boundaries
In AI systems, pay special attention to these boundaries.
Boundary 1 — User → AI
Controls:
Authentication
Input validation
Rate limiting
Boundary 2 — AI → Data
Controls:
Authorization
Tenant isolation
DLP
Retrieval filtering
Boundary 3 — AI → Action
Controls:
Tool authorization
Parameter validation
Policy
Human approval
Audit
Visualized:
User
│
▼
[ AI SYSTEM ]
│
├──────────→ Data
│
└──────────→ Actions
Each arrow should have explicit controls.
67. Architecture Principle: Least Privilege
This applies everywhere.
Not just users.
Also:
Agents
Tools
Services
Models
Workers
Databases
Caches
For example:
Security Triage Agent
│
├── READ Jira
├── READ GitHub
├── READ security docs
└── CREATE Jira comment
NOT:
├── DELETE Jira
├── WRITE production DB
└── READ secrets
68. Architecture Principle: Fail Closed for High-Risk Actions
Suppose an authorization service becomes unavailable.
For a harmless read:
Authorization unavailable
↓
Maybe retry
For a destructive production action:
Authorization unavailable
↓
DO NOT EXECUTE
This is a classic security architecture principle:
When the cost of an incorrect authorization decision is high, failure should not silently become permission.
69. Architecture Principle: Make AI Decisions Observable
You don't necessarily need to expose internal model reasoning to users.
But you should make system behavior auditable.
For example:
Agent Task: SEC-4321
Tools used:
1. Jira
2. GitHub
3. Scanner
Policy checks:
3 passed
1 blocked
Final action:
Jira comment created
This gives operators visibility without requiring the system to expose private model reasoning.
70. Architecture Interview Question: "Why Did You Choose This Architecture?"
A strong answer should sound like:
"I started from workload characteristics rather than technology preferences. The application requires enterprise knowledge retrieval, occasional multi-step actions, strong tenant isolation, and auditable operations. Therefore I separated deterministic application controls from probabilistic AI reasoning, used a dedicated retrieval layer, introduced a controlled tool gateway for agents, and centralized identity, policy, observability, and cost controls."
That's architecture reasoning.
Not:
"We used Kubernetes because Kubernetes is scalable."
71. Architecture Interview Question: "How Do You Design for Scale?"
A strong answer:
"I first identify the bottleneck rather than scaling every component equally. I separate stateless APIs from durable state, use queues for asynchronous workloads, scale workers horizontally, cache expensive operations where safe, control downstream model and tool quotas, and monitor end-to-end latency and throughput."
72. Architecture Interview Question: "How Do You Control Cost?"
A strong answer:
"I optimize the entire request path: retrieval quality, context size, model routing, caching, agent step limits, batching, and token budgets. I measure cost per request and per workflow rather than looking only at the monthly provider bill."
73. Architecture Interview Question: "How Do You Secure an AI Platform?"
A strong answer:
"I treat the LLM as an untrusted probabilistic component rather than a security boundary. Identity and authorization are enforced outside the model, tools are least-privileged and policy-controlled, tenant isolation is enforced throughout retrieval and storage, sensitive data is controlled across the lifecycle, and agent actions are audited and bounded."
74. Architecture Interview Question: "How Do You Handle a Model Outage?"
A good answer should distinguish between workloads.
For a low-risk chatbot:
Primary model
↓
Fallback model
↓
Response
For a high-risk security workflow:
Primary model unavailable
↓
Fallback unavailable
↓
Safe failure
↓
Human escalation
The appropriate fallback depends on the risk.
75. Architecture Interview Question: "How Do You Move From POC to Production?"
A strong answer:
"I revisit the design across six dimensions: security, reliability, scale, cost, observability, and governance. I replace local state with durable services where needed, introduce authentication and authorization, productionize ingestion and indexing, add monitoring and evaluation, establish CI/CD and versioning, define failure and recovery behavior, and validate the system under realistic workloads."
76. A Practical Example — Production Security Copilot
Let's imagine we are building:
Enterprise Security Copilot
Capabilities:
search security standards
inspect Jira vulnerabilities
inspect GitHub pull requests
run approved security checks
summarize evidence
create a Jira comment
generate a security report
Architecture:
Security Engineer
│
▼
Security Copilot
│
┌────────┼─────────┐
▼ ▼ ▼
RAG Agent Memory
│ │
│ ▼
│ Tool Gateway
│ │
│ ┌────┼────┐
│ ▼ ▼ ▼
│ Jira GitHub Scanner
│
▼
Security KB
│
Vector + Search
Security controls:
Identity
↓
Tenant / RBAC
↓
Tool Authorization
↓
Policy Engine
↓
Approval for sensitive actions
↓
Audit
Operational controls:
Tracing
Metrics
Cost controls
Token budgets
Agent step limits
Retries
Timeouts
Now this is an enterprise architecture rather than a chatbot.
77. A Production Readiness Checklist
Before calling an AI application "production ready", ask:
Architecture
□ Clear component boundaries
□ Appropriate data stores
□ Stateless scaling where possible
□ Durable state where required
□ Failure modes defined
AI
□ Model selected using evaluation
□ Prompt/version management
□ RAG evaluation
□ Agent evaluation
□ Model upgrade process
Security
□ Authentication
□ Authorization
□ Tenant isolation
□ Secrets management
□ DLP / sensitive data controls
□ Tool security
□ Prompt injection testing
□ Audit
Operations
□ Metrics
□ Tracing
□ Logging
□ Alerting
□ Cost monitoring
□ Rate limiting
□ Capacity planning
Resilience
□ Timeouts
□ Retry policy
□ Circuit breakers
□ Queueing where needed
□ Backups
□ Disaster recovery
78. The Architecture Principles I Would Remember
There are hundreds of individual design decisions.
But these principles capture most of them.
Principle 1
Don't make the LLM your security boundary.
Principle 2
Separate reasoning, execution, and authorization.
Principle 3
Treat retrieval as a production data pipeline.
Principle 4
Treat prompts, models, indexes, and policies as versioned artifacts.
Principle 5
Give agents narrow capabilities instead of unrestricted power.
Principle 6
Every autonomous workflow needs limits.
Principle 7
Optimize cost and latency across the complete request path.
Principle 8
Monitor quality, not just infrastructure health.
Principle 9
Design for failure before designing for scale.
Principle 10
Use deterministic controls around probabilistic components.
79. The Production AI Mental Model
At this point, the entire series can be summarized as:
ENTERPRISE AI
USER
│
▼
APPLICATION
│
┌───────────┴───────────┐
▼ ▼
RAG AGENT
│ │
KNOWLEDGE ACTION
│ │
└───────────┬───────────┘
▼
LLM
│
▼
┌───────────────────────┐
│ Deterministic Controls│
│ │
│ IAM │
│ Authorization │
│ Policy │
│ Validation │
│ Rate Limits │
│ Cost Controls │
│ Audit │
└──────────┬────────────┘
│
▼
ENTERPRISE SYSTEMS
The model provides intelligence.
The surrounding architecture provides:
Control
Security
Reliability
Scale
Governance
That distinction is fundamental.
80. Final Takeaways
A production AI architecture is not:
LLM
+
Vector DB
+
Chat UI
It is a complete distributed system.
You need to reason about:
Production AI
│
┌───────────────┼────────────────┐
▼ ▼ ▼
Intelligence Data Execution
│ │ │
LLM RAG Tools
│ │ │
└───────────────┼────────────────┘
▼
Platform
│
┌───────────────┼─────────────────┐
▼ ▼ ▼
Security Observability Reliability
│ │ │
└───────────────┼─────────────────┘
▼
Scale
│
▼
Cost
The most important architectural lesson is:
A production AI system is a distributed application with an AI component—not an AI model surrounded by a few APIs.
And one final principle ties the entire series together:
Use AI where it provides intelligence; use deterministic software where you need guarantees.
Let the model reason.
Let retrieval provide knowledge.
Let agents coordinate actions.
But let identity, authorization, policies, validation, transactions, and security controls remain deterministic and enforceable outside the model.
What's Next?
We've now covered the architecture of:
LLMs
↓
RAG
↓
Agents
↓
Production AI Platform
The obvious next question is the one that matters most to security engineers:
"What can actually go wrong?"
A production AI system introduces a new set of attack surfaces:
User
↓
Prompt
↓
LLM
↓
RAG
↓
Memory
↓
Tools
↓
Enterprise Systems
An attacker may try to manipulate:
Prompts
Documents
Embeddings
Memory
Tool calls
Model outputs
And an innocent failure can be just as dangerous as an intentional attack when an autonomous agent has access to real systems.
That brings us to Part 5 — AI Security: Threat Modeling and Securing LLM, RAG & Agentic AI Applications, where we'll take the architecture above and attack it from a security-engineering perspective.
We'll cover:
Prompt injection and indirect prompt injection
RAG poisoning
Sensitive information disclosure
Cross-tenant data leakage
Excessive agency
Agent/tool abuse
Output injection
Memory poisoning
AI supply-chain security
Model and plugin trust
SSRF and network risks in AI agents
Authorization architecture
Human-in-the-loop security
AI red teaming
Security test cases
AI threat modeling
Secure reference architecture
A complete enterprise AI security checklist

