Skip to main content

Command Palette

Search for a command to run...

AI Engineering Interview Guide — Part 4: Production AI Architecture

Updated
•34 min read•View as Markdown
S
I like breaking things that are supposed to be secure. When I’m not hunting vulnerabilities, I’m exploring systems, architectures, and the assumptions behind them.

How to turn an LLM/RAG/Agent POC into a secure, scalable, observable, cost-efficient enterprise AI platform

In the previous three articles, we built the conceptual foundation:

Part 1 — LLM Fundamentals

We learned about tokens, embeddings, vectors, context windows, model parameters, and fine-tuning.

Part 2 — RAG Deep Dive

We explored ingestion, chunking, embeddings, retrieval, reranking, metadata, evaluation, and RAG security.

Part 3 — Agentic AI

We looked at agents, tools, planning, memory, orchestration, autonomy, and agent security.

Now comes the question that separates a prototype from a real production system:

"How would you architect this for an enterprise?"

A POC can often be built with:

Python
   +
LLM API
   +
Simple vector store
   +
A notebook

Production is very different.

Now you need to think about:

Security
Scalability
Availability
Latency
Cost
Data governance
Multi-tenancy
Observability
Failure handling
Versioning
Disaster recovery
Compliance

And the architecture becomes something more like:

                              USERS
                                │
                                ▼
                         ┌─────────────┐
                         │ API / Web   │
                         │ Application │
                         └──────┬──────┘
                                │
                                ▼
                        Identity / IAM
                                │
                                ▼
                       AI Gateway Layer
                                │
                 ┌──────────────┼──────────────┐
                 ▼              ▼              ▼
                LLM            RAG           Agents
                 │              │              │
                 │         ┌────┴────┐     Tool Gateway
                 │         ▼         ▼          │
                 │      Search    Vector DB     │
                 │                              │
                 └──────────────┬───────────────┘
                                ▼
                     Enterprise Data / APIs
                                │
                                ▼
                      Observability / Audit

This article focuses on architecture decisions, not framework tutorials.


1. The First Architecture Question: What Are We Actually Building?

One of the biggest mistakes in AI architecture is starting with:

"Which LLM should we use?"

That's too early.

Start with the workload.

For example:

Business Requirement
        │
        ▼
What does the system actually do?
        │
        ├── Generate text?
        ├── Answer questions?
        ├── Search enterprise data?
        ├── Make decisions?
        ├── Execute actions?
        └── Perform multi-step tasks?

These map to different architectures.

Simple generation

User → LLM → Response

Enterprise knowledge assistant

User → RAG → LLM

Autonomous task assistant

User → Agent → Tools → Enterprise Systems

Complex enterprise AI platform

User
 ↓
AI Application
 ↓
RAG + Agents + APIs + Enterprise Data

The first architectural principle is therefore:

Design around the workload and risk, not around the model.


2. POC Architecture vs Production Architecture

Let's take a practical example.

Imagine an engineer creates a security assistant POC:

User
 ↓
Python script
 ↓
LLM API
 ↓
FAISS
 ↓
PDF files

It works.

The demo looks great.

Then someone says:

"Let's make this available to 20,000 employees."

The architecture changes immediately.

You now need:

                   USERS
                     │
                     ▼
                 API Gateway
                     │
                Authentication
                     │
                     ▼
              Application Layer
                     │
          ┌──────────┼──────────┐
          ▼          ▼          ▼
         RAG        Agent       LLM
          │          │
          ▼          ▼
       Search       Tools
          │          │
          ▼          ▼
        Data      Enterprise APIs

And around it:

IAM
Secrets
Observability
Rate Limiting
Caching
Audit
Security Controls
CI/CD

A useful interview answer is:

"My POC architecture optimizes for learning and speed. Production architecture optimizes for security, reliability, scale, cost, operability, and governance."


3. A Reference Enterprise AI Architecture

Let's build a reusable reference architecture.

                              ┌───────────────┐
                              │     Users     │
                              └───────┬───────┘
                                      │
                                      ▼
                            ┌──────────────────┐
                            │ API / Web Layer  │
                            └────────┬─────────┘
                                     │
                                     ▼
                          ┌────────────────────┐
                          │ Identity / Access  │
                          └─────────┬──────────┘
                                    │
                                    ▼
                         ┌─────────────────────┐
                         │   AI Gateway /      │
                         │   Application       │
                         └─────────┬───────────┘
                                   │
                ┌──────────────────┼──────────────────┐
                ▼                  ▼                  ▼
          ┌──────────┐       ┌──────────┐       ┌───────────┐
          │   LLM    │       │   RAG    │       │   Agent   │
          └──────────┘       └────┬─────┘       └─────┬─────┘
                                  │                    │
                           ┌──────┴──────┐         Tool Gateway
                           ▼             ▼              │
                       Search       Vector Store        │
                                                        │
                                    ┌───────────────────┼────────────────┐
                                    ▼                   ▼                ▼
                                  Jira               GitHub             DB

Supporting services:

Secrets Management
Policy / Authorization
Cache
Observability
Audit Logging
DLP
Rate Limiting
Configuration
Evaluation

The exact technologies will change from company to company.

The architectural responsibilities should not.


4. Separate the AI Layer From the Application Layer

An enterprise AI system should not become:

Frontend
 ↓
LLM
 ↓
Everything

Instead, separate responsibilities.

                    Application
                        │
                        ▼
                  AI Orchestration
                        │
         ┌──────────────┼──────────────┐
         ▼              ▼              ▼
        LLM            RAG           Tools

Why?

Because the application must control:

  • authentication

  • authorization

  • business rules

  • validation

  • transactions

  • rate limits

  • logging

  • error handling

The LLM should not own these responsibilities.

A useful architectural principle is:

Keep probabilistic reasoning isolated from deterministic system controls.


5. Introduce an AI Gateway

If many applications directly call different model providers:

App A → Provider A
App B → Provider B
App C → Provider C
App D → Provider A

model management becomes difficult.

A common enterprise pattern is:

Applications
     │
     ▼
 AI Gateway
     │
 ┌───┼───────────────┐
 ▼   ▼               ▼
LLM A LLM B       Local Model

The gateway can centralize:

  • model routing

  • authentication

  • quotas

  • rate limiting

  • logging

  • policy enforcement

  • cost tracking

  • provider abstraction

  • model version management

This is especially useful when the organization uses multiple models.


6. Why Model Abstraction Matters

Imagine your application directly depends on a provider-specific API.

Later:

The provider changes the model.

or:

A cheaper model becomes available.

or:

The organization wants an on-prem model.

If every application is tightly coupled to one provider, migration becomes painful.

A model abstraction layer can provide:

Application
     │
     ▼
Model Interface
     │
 ┌───┼──────────────┐
 ▼   ▼              ▼
Model A Model B   Local Model

This doesn't mean every model should behave identically.

It means the application should avoid unnecessary provider coupling.


7. Model Selection Is an Architecture Decision

Don't ask:

"Which model is the best?"

Ask:

"Which model is appropriate for this workload?"

Consider:

Quality
Latency
Cost
Context size
Reasoning capability
Multilingual support
Privacy requirements
Data residency
Availability
Rate limits
Operational maturity

For example:

Simple classification
        ↓
Smaller model

Complex reasoning
        ↓
Larger model

Sensitive workload
        ↓
Approved private deployment

This naturally leads to model routing.


8. Model Routing

Suppose an enterprise platform receives 1 million AI requests every day.

Not all requests are equally complex.

A router can classify the request:

                      Request
                         │
                         ▼
                   Task Router
                         │
              ┌──────────┼──────────┐
              ▼          ▼          ▼
           Simple      Medium     Complex
              │          │          │
              ▼          ▼          ▼
           Model A     Model B     Model C

This can reduce cost while maintaining quality.

But the router itself should be measured.

Don't assume:

"Smaller model = always good enough."

Use evaluation data to validate the trade-off.


9. Data Architecture

AI systems often touch many different kinds of data.

For example:

Documents
Structured records
Embeddings
User conversations
Agent state
Application data
Logs
Telemetry
Model outputs
Evaluation data

Trying to put everything into one database is usually a mistake.

A better design might be:

                   Data Layer
                       │
       ┌───────────────┼────────────────┐
       ▼               ▼                ▼
 Object Storage    Relational DB    Vector Store
       │               │                │
 Documents         Business Data     Embeddings

And possibly:

Cache
Search Index
Graph Store
Event Store

Choose the storage technology based on the workload.


10. Object Storage vs Database vs Vector Store

A simple rule of thumb:

Object storage

Good for:

  • PDFs

  • documents

  • large files

  • raw datasets

Relational database

Good for:

  • transactions

  • users

  • metadata

  • configuration

  • business entities

Vector store

Good for:

  • embeddings

  • semantic retrieval

Search engine

Good for:

  • keyword retrieval

  • filtering

  • analytics-style search

Cache

Good for:

  • frequently accessed data

  • expensive computations

  • low-latency responses

The architecture should reflect the data's access pattern.


11. RAG in Production Is a Data Pipeline

A production RAG architecture isn't simply:

PDF → Vector DB

It is more like:

                    Source Systems
                         │
                         ▼
                    Ingestion
                         │
                         ▼
               Parse / Normalize
                         │
                         ▼
                    Classify
                         │
                         ▼
                    Chunk
                         │
                         ▼
                  Metadata
                         │
                         ▼
                    Embed
                         │
                         ▼
                 Index / Publish
                         │
                         ▼
                  Search Layer

This should behave like a proper production data pipeline.

You need:

  • retries

  • dead-letter handling

  • idempotency

  • monitoring

  • versioning

  • change detection

  • delete propagation


12. Event-Driven Document Ingestion

Instead of periodically scanning everything:

Every night
   ↓
Scan all documents

you can use events:

Document Updated
       │
       ▼
     Event
       │
       ▼
 Ingestion Worker
       │
       ▼
Parse → Chunk → Embed → Index

This makes updates more responsive and can reduce unnecessary work.

The same pattern applies to deletion:

Document Deleted
       ↓
Deletion Event
       ↓
Remove corresponding chunks
       ↓
Invalidate cache/index

13. Idempotency Is Important in AI Pipelines

Suppose the document ingestion worker fails halfway through.

A retry happens.

Without idempotency you might create:

Document A
 ├── Chunk 1
 ├── Chunk 2
 ├── Chunk 3
 ├── Chunk 1   ← duplicate
 ├── Chunk 2   ← duplicate
 └── Chunk 3   ← duplicate

Use stable identifiers.

For example:

document_id
document_version
chunk_id

so the pipeline can safely retry.

AI systems inherit many of the same distributed-systems problems as conventional applications.


14. Multi-Tenancy Must Be Designed, Not Added Later

Suppose your platform supports:

Tenant A
Tenant B
Tenant C

You need a clear answer to:

Where does tenant isolation happen?

Potential boundaries include:

Tenant
 ↓
Application
 ↓
Database
 ↓
Vector Index
 ↓
Cache
 ↓
Logs

One common logical model is:

tenant_id = A

being propagated through the entire request.

User
 ↓
Tenant Context
 ↓
RAG Retrieval
 ↓
Tools
 ↓
Database

This becomes extremely important because AI systems often combine information from multiple sources.


15. Cache Design Can Create Security Problems

Caching is useful.

For example:

User Query
 ↓
Cache
 ↓
Cached response

But imagine:

Tenant A asks:
"What is our architecture?"

and the response is cached using only:

query = "What is our architecture?"

Then Tenant B asks the same question.

Oops.

The cached answer could be returned to the wrong tenant.

The cache key may need to include relevant security context:

tenant_id
user_scope
authorization context
query

This is a good example of a broader principle:

Every optimization layer must preserve the security model.


16. Authorization Must Flow Through the Architecture

Suppose a user asks an agent:

"Show me all security tickets."

The request may travel through:

User
 ↓
Application
 ↓
Agent
 ↓
Jira Tool
 ↓
Jira

The authorization context must not disappear along the way.

Conceptually:

User Identity
     │
     ▼
Authorization Context
     │
     ▼
Agent
     │
     ▼
Tool Gateway
     │
     ▼
Enterprise Resource

The agent may be autonomous, but it still operates within a user's security context.


17. Don't Let Business Rules Live in the Prompt

Imagine a business rule:

"Only security managers can approve production exceptions."

Don't implement this as:

System Prompt:

If user is not a security manager,
do not approve.

That is useful as guidance but weak as enforcement.

Instead:

Agent
 ↓
Approval Request
 ↓
Authorization Service
 ↓
Policy
 ↓
ALLOW / DENY

The deterministic policy system owns the business rule.


18. Human Approval Architecture

For sensitive actions, introduce an approval boundary.

Example:

                 Agent
                   │
                   ▼
             Proposed Action
                   │
                   ▼
              Risk Engine
                   │
            ┌──────┴──────┐
            ▼             ▼
          Low Risk      High Risk
            │             │
            ▼             ▼
         Execute      Human Approval
                            │
                       ┌────┴────┐
                       ▼         ▼
                    Approve    Reject
                       │
                       ▼
                    Execute

This is especially useful for:

  • production changes

  • financial operations

  • destructive actions

  • external communication

  • privileged access changes


19. Design for Failure

Production AI systems will fail.

A model might be unavailable.

A vector store might be unavailable.

A tool may time out.

A provider may throttle requests.

An architecture should explicitly define what happens.

               Request
                  │
                  ▼
                Agent
                  │
                  ▼
                Tool
                  │
               Failure
                  │
        ┌─────────┼─────────┐
        ▼         ▼         ▼
      Retry    Fallback   Escalate

But retries require care.


20. Retry Isn't Always Safe

Consider:

create_payment()

The request times out.

You don't know whether the payment succeeded.

Retrying may create a duplicate transaction.

Therefore:

Retries need idempotency semantics, especially for write operations.

A useful architecture includes:

Request ID
Idempotency Key
Timeout
Retry Policy
Circuit Breaker

This is a distributed-systems lesson that applies directly to agentic AI.


21. Timeouts and Cancellation

Imagine an agent starts:

Tool A → 10 sec
Tool B → 15 sec
Tool C → 30 sec
Tool D → hangs

Without controls, the user may wait indefinitely.

Set:

Request timeout
Tool timeout
LLM timeout
Maximum agent duration

And make cancellation possible.

For an agent system:

Cancel Request
     ↓
Orchestrator
     ↓
Stop pending tool calls
     ↓
Terminate execution

22. Agent Budgets

Agents can consume unpredictable resources.

Set explicit limits:

Maximum steps
Maximum tool calls
Maximum tokens
Maximum execution time
Maximum estimated cost

For example:

MAX_STEPS       = 15
MAX_TOOL_CALLS  = 25
MAX_RUNTIME     = 2 minutes

The numbers are workload-specific.

The architectural principle is what matters:

Every autonomous loop should have a budget.


23. Observability Is Not Optional

In a traditional API, you might log:

Request
Response
Latency
Status

For an agentic system, that's not enough.

You may need to understand:

User request
 ↓
LLM call
 ↓
Tool selection
 ↓
Tool parameters
 ↓
Tool result
 ↓
Next LLM call
 ↓
Another tool
 ↓
Final answer

This is why distributed tracing becomes especially valuable.


24. Trace the Entire AI Request

A useful trace might look like:

Request ID: 12345

├── Authentication
├── Query processing
├── Embedding generation
├── Retrieval
│    ├── Vector search
│    └── Reranking
├── LLM call #1
│    └── Tool: get_jira_ticket
├── Jira API
├── LLM call #2
│    └── Tool: get_github_pr
├── GitHub API
├── LLM call #3
└── Final response

Now an engineer can answer:

Why did this request take 18 seconds?

or:

Why did the agent call Jira four times?


25. What Should You Measure?

A mature AI platform has multiple dimensions of observability.

Infrastructure

CPU
Memory
Network
Storage

AI

Token usage
Latency
Model errors
Context size

RAG

Retrieval latency
Recall@K
No-result rate
Reranking latency

Agent

Steps
Tool calls
Tool failures
Task completion

Business

Success rate
User satisfaction
Escalation rate

Security

Policy violations
Blocked actions
Sensitive-data detections
Unauthorized access attempts

26. Cost Architecture

AI systems can become expensive surprisingly quickly.

Consider:

1 million requests/day
×
5,000 input tokens
×
large model

That can become a significant operational cost.

Cost should therefore be designed into the architecture.

                    Cost Control
                         │
          ┌──────────────┼──────────────┐
          ▼              ▼              ▼
       Caching       Model Routing    Retrieval
          │              │              │
          ▼              ▼              ▼
       Fewer calls     Cheaper models   Less context

27. Token Economics

Consider an agent workflow:

Step 1 → 3,000 tokens
Step 2 → 4,000 tokens
Step 3 → 5,000 tokens
Step 4 → 4,000 tokens
Step 5 → 6,000 tokens

The system may process tens of thousands of tokens for one task.

Cost optimization therefore isn't only:

"Use a cheaper model."

You should also ask:

Why are we sending so much context?

Potential improvements:

  • better retrieval

  • shorter prompts

  • context compression

  • caching

  • fewer agent loops

  • smaller models for simple steps

  • structured outputs


28. Caching Strategy

Different things can be cached.

Embeddings

Avoid generating the same embedding repeatedly.

Retrieval results

Useful when queries repeat and the underlying data hasn't changed.

LLM responses

Useful for deterministic or low-volatility workloads.

Tool results

Potentially useful for safe, short-lived data.

But every cache needs to respect:

Freshness
Authorization
Tenant isolation
Data sensitivity

29. Semantic Caching

Traditional caching depends on exact keys.

"How do I reset my password?"

Semantic caching can recognize:

"How can I change my forgotten password?"

as potentially similar.

Conceptually:

New Query
   ↓
Semantic Similarity
   ↓
Existing Cached Query
   ↓
Reuse result?

This can reduce model calls.

But semantic caching is risky for sensitive workloads if the cache scope and authorization model aren't designed carefully.


30. Latency Architecture

An AI request might involve:

Authentication
+
Embedding
+
Retrieval
+
Reranking
+
LLM
+
Tool Calls

Latency can accumulate.

A simple mental model is:

Total Latency
≈
Network
+
Retrieval
+
Reranking
+
LLM
+
Tool Calls

This is why architecture needs to distinguish between:

  • sequential calls

  • parallel calls

  • asynchronous work

  • streaming


31. Parallel Tool Execution

Suppose an agent needs:

Jira ticket
GitHub PR
Security policy

If these are independent:

Agent
 ├── Jira
 ├── GitHub
 └── RAG

they may be executed in parallel rather than:

Jira → GitHub → RAG

Parallel execution can reduce latency.

But only when:

  • operations are independent

  • resource limits are respected

  • authorization is enforced consistently


32. Streaming Responses

Users don't necessarily need to wait for the entire response.

Instead:

LLM
 ↓
Token stream
 ↓
User sees partial response

This can improve perceived latency.

For agentic systems, however, don't confuse:

streaming intermediate text

with:

streaming dangerous actions.

Actions should remain governed and auditable.


33. Asynchronous Architecture

Some AI tasks don't need synchronous responses.

For example:

"Analyze 10,000 security findings and generate a report."

Instead of keeping an HTTP request open:

User
 ↓
Submit Job
 ↓
Queue
 ↓
Worker
 ↓
Agent / LLM
 ↓
Store Result
 ↓
Notify User

This is often much more resilient.


34. Queue-Based Architecture

For long-running AI jobs:

                API
                 │
                 ▼
               Queue
                 │
       ┌─────────┼─────────┐
       ▼         ▼         ▼
    Worker A  Worker B  Worker C
       │         │         │
       └─────────┼─────────┘
                 ▼
             AI Services

Benefits include:

  • load smoothing

  • retry management

  • backpressure

  • asynchronous processing

  • worker scaling


35. Backpressure

Imagine:

10 requests/sec

suddenly becomes:

10,000 requests/sec

Your LLM provider may not handle that load.

A queue allows the application to absorb bursts.

Users
 ↓
API
 ↓
Queue
 ↓
Controlled workers
 ↓
LLM

This protects downstream systems.


36. Rate Limiting

Rate limiting should exist at multiple levels.

User
 ↓
Tenant
 ↓
Application
 ↓
Model
 ↓
Tool

For example:

User:
100 requests/hour

Tenant:
10,000 requests/hour

Tool:
500 calls/minute

These are examples only.

The correct values depend on the business requirements.


37. Disaster Recovery

What happens if your AI platform loses:

Vector DB
Database
Object storage
Configuration
Agent state

You need recovery plans.

For critical systems:

Primary
   │
   ├── Backups
   ├── Replication
   └── Recovery procedures

For RAG specifically, remember:

The source documents are the authoritative data. Embeddings and indexes can often be rebuilt from them.

This can simplify disaster recovery planning.


38. Rebuild vs Backup

This is an interesting architecture decision.

Suppose the vector index is destroyed.

If you have:

Source documents
+
Chunking configuration
+
Embedding model/version
+
Metadata

you may be able to reconstruct it.

Therefore, the question becomes:

Is the vector index a primary data store or a derived index?

In many architectures:

Authoritative Source
        ↓
     Derived
        ↓
Embedding / Index

This distinction can simplify recovery.


39. Version Everything

AI systems change frequently.

You should consider versioning:

Model
Prompt
Embedding model
Chunking strategy
Reranker
Knowledge index
Tool schemas
Agent policy
Evaluation dataset

For example:

Model:            model-v4
Prompt:           prompt-v17
Embedding:        embed-v3
Index:            index-2026-09
Policy:           policy-v8

Now when quality changes, you can investigate what changed.


40. Prompt Versioning Is Engineering

Prompts should not live only in someone's laptop.

Think of prompts as software artifacts.

Prompt
 ↓
Git
 ↓
Review
 ↓
Test
 ↓
Deploy
 ↓
Monitor

A production prompt change should be treated similarly to code changes.


41. Evaluation Before Deployment

Suppose you want to replace:

Model A

with:

Model B

Don't simply deploy it.

Run an evaluation set:

                   Evaluation Dataset
                           │
                    ┌──────┴──────┐
                    ▼             ▼
                  Model A       Model B
                    │             │
                    └──────┬──────┘
                           ▼
                       Compare

Compare:

  • correctness

  • groundedness

  • latency

  • cost

  • safety

  • tool behavior

This is especially important because model changes can affect security behavior too.


42. Shadow Testing

A useful production technique is to send the same request to a new model without using its result operationally.

User Request
     │
     ├── Production Model → User
     │
     └── Candidate Model → Evaluation Only

Then compare results.

This reduces migration risk.


43. Canary Releases

Instead of moving all users immediately:

100% Model A

you can introduce:

95% → Model A
5%  → Model B

Monitor:

Error rate
Latency
Cost
Quality
Security events

Then gradually increase the percentage.

The same principle works for:

  • prompts

  • retrieval pipelines

  • agent policies

  • tool versions


44. AI CI/CD

A mature AI platform should treat AI changes as deployable artifacts.

A simplified pipeline:

Code / Prompt / Model Change
           │
           ▼
       Unit Tests
           │
           ▼
      AI Evaluation
           │
           ▼
   Security Test Suite
           │
           ▼
      Integration Tests
           │
           ▼
       Canary Deploy
           │
           ▼
       Production

This is essentially CI/CD for AI systems.


45. Traditional DevSecOps Still Matters

AI does not eliminate traditional security engineering.

The application still has:

APIs
Containers
Kubernetes
Databases
Identity
Secrets
Dependencies
Networks
Cloud infrastructure

Therefore, you still need:

SAST
DAST
SCA
Secrets scanning
Container scanning
API security
IAM review
Infrastructure security

And add AI-specific testing:

Prompt injection
RAG poisoning
Agent tool abuse
Data leakage
Excessive agency

The result is:

Traditional AppSec
       +
AI Security

46. Secrets Management

Never put:

LLM_API_KEY=...

inside:

source code
prompt
agent memory
RAG documents
logs

Use dedicated secret-management mechanisms.

Conceptually:

Application
     │
     ▼
Secret Manager
     │
     ▼
Temporary Credential
     │
     ▼
LLM / Tool

Rotate secrets and limit their scope.


47. Network Architecture

An enterprise AI platform may contain:

Internet-facing API
Internal RAG services
Vector database
Model endpoints
Tool servers
Databases

These shouldn't all live in the same unrestricted network.

A simplified model:

                Internet
                    │
                    ▼
               API Gateway
                    │
               Public Zone
                    │
                    ▼
              AI Application
                    │
        ┌───────────┼───────────┐
        ▼           ▼           ▼
      RAG        Tool Layer    LLM
        │           │
        ▼           ▼
   Data Layer   Enterprise APIs

Apply network segmentation and restricted egress where appropriate.


48. Observability Data Is Sensitive

There is a trap here.

An engineer might say:

"We'll log every prompt and every response for debugging."

Sounds useful.

But those logs may contain:

PII
Secrets
Customer data
Source code
Confidential documents

So AI observability requires:

Redaction
Classification
Access Control
Retention
Encryption

The debugging system itself becomes part of the security architecture.


49. Data Retention

Ask:

How long should the system retain?

Potentially:

Prompt
Response
Conversation
Agent trace
Retrieved documents
Tool results
Evaluation results

You may not want all of them retained forever.

A production design should define:

Purpose
Retention Period
Access
Deletion
Legal / Compliance Requirements

50. Model and Data Residency

For enterprise workloads, ask:

Where does the data physically or logically go?

A request might travel:

India
 ↓
Application
 ↓
US Model Provider
 ↓
US Logs

That may or may not be acceptable depending on organizational requirements.

Model architecture should therefore consider:

  • geographic residency

  • provider data handling

  • contractual controls

  • encryption

  • private endpoints

  • local models

This is an architecture requirement, not merely a compliance checkbox.


51. Open-Source Models vs Hosted Models

Both have trade-offs.

Hosted model

Advantages:

Fast adoption
Managed infrastructure
Less operational burden

Trade-offs may include:

Provider dependency
Data-handling considerations
External availability
Cost

Self-hosted model

Advantages:

More control
Potentially stronger isolation
Customization

Trade-offs:

GPU infrastructure
Operations
Scaling
Patching
Model lifecycle
Cost

Don't ask:

"Which is better?"

Ask:

"Which operational model fits our security, cost, performance, and data requirements?"


52. Kubernetes Architecture

For organizations already operating Kubernetes, an AI platform may look like:

                 Kubernetes Cluster
                        │
       ┌────────────────┼────────────────┐
       ▼                ▼                ▼
 AI Gateway          RAG Service      Agent Service
       │                │                │
       ▼                ▼                ▼
    Model GW         Vector DB       Tool Gateway

But don't assume everything must run inside the cluster.

For example:

Kubernetes
    │
    ├── Application
    ├── RAG
    └── Agent
          │
          ▼
    Managed Model Service

Architecture should follow operational requirements.


53. Scaling Strategy

Different components scale differently.

Component          Scaling Pattern

API                 Horizontal
RAG workers         Horizontal
Embeddings          Batch / worker scaling
LLM                 Provider-dependent
Vector DB           Index/compute scaling
Agents              Worker scaling
Tools               Per-service scaling

Don't scale everything identically.

The bottleneck may be different at each layer.


54. A Common Architecture Mistake: Scaling the Wrong Layer

Suppose:

API servers = 100

but:

LLM quota = exhausted

Adding more API instances doesn't solve the actual bottleneck.

Likewise:

10 agent workers

may be useless if:

Tool API = 50 requests/minute

Production architecture requires capacity planning across the entire dependency chain.


55. Reliability Budget

A useful architectural mindset is to think in terms of:

Availability
Latency
Cost
Quality
Security

These often conflict.

For example:

More retrieval
   ↓
Better recall
   ↓
More tokens
   ↓
Higher cost
   ↓
Higher latency

Or:

Bigger model
   ↓
Potentially better quality
   ↓
Higher cost
   ↓
Higher latency

Architecture is therefore about managing trade-offs.


56. AI Quality Is a Production Concern

Traditional monitoring might say:

HTTP 200

But the AI answer could still be terrible.

Therefore, production monitoring should include quality signals.

For RAG:

Groundedness
Retrieval quality
No-answer rate

For agents:

Task completion
Tool accuracy
Unnecessary steps
Policy violations

This is one of the biggest differences between AI systems and traditional APIs.


57. AI SRE Mindset

A production AI team should ask:

"Is the system healthy?"

not simply:

"Is the HTTP endpoint up?"

Health should include:

Infrastructure
+
Model availability
+
Retrieval quality
+
Agent behavior
+
Latency
+
Cost
+
Security

You can have:

99.99% API uptime

and still have a completely unusable AI system because retrieval quality collapsed.


58. A Production AI Incident Example

Imagine users report:

"The assistant suddenly started giving incorrect security answers."

Infrastructure is healthy.

API is healthy.

Model endpoint is healthy.

What changed?

Yesterday:
Embedding Model v2

Today:
Embedding Model v3

The new embedding model changed the retrieval distribution.

Now:

Recall@5
92% → 61%

This is why AI observability must connect:

Model Version
+
Embedding Version
+
Index Version
+
Prompt Version
+
Quality Metrics

Without those links, debugging becomes guesswork.


59. Architecture for Safe Model Upgrades

A robust process is:

New Model
   │
   ▼
Offline Evaluation
   │
   ▼
Security Evaluation
   │
   ▼
Shadow Traffic
   │
   ▼
Canary
   │
   ▼
Production

This is much safer than:

New model → deploy everywhere

60. A Full Production AI Platform

Let's bring everything together.

                              USERS
                                │
                                ▼
                        ┌────────────────┐
                        │ API / Web App  │
                        └───────┬────────┘
                                │
                                ▼
                       ┌──────────────────┐
                       │ Identity / IAM   │
                       └────────┬─────────┘
                                │
                                ▼
                     ┌─────────────────────┐
                     │    AI Gateway       │
                     │                     │
                     │ Routing             │
                     │ Rate Limits         │
                     │ Policy              │
                     │ Cost Controls       │
                     └──────────┬──────────┘
                                │
           ┌────────────────────┼─────────────────────┐
           ▼                    ▼                     ▼
       ┌────────┐          ┌────────┐          ┌──────────┐
       │  LLM   │          │  RAG   │          │  Agent   │
       └────────┘          └───┬────┘          └────┬─────┘
                               │                    │
                         ┌─────┴─────┐         Tool Gateway
                         ▼           ▼              │
                      Search      Vector DB         │
                                                    │
                                      ┌─────────────┼─────────────┐
                                      ▼             ▼             ▼
                                    Jira         GitHub          DB

                 ┌────────────────────────────────────────┐
                 │      Data / Platform Layer              │
                 │                                        │
                 │ Object Storage | SQL | Cache | Events  │
                 └────────────────────────────────────────┘

                 ┌────────────────────────────────────────┐
                 │      Governance / Security              │
                 │                                        │
                 │ IAM | DLP | Secrets | Audit | Policy  │
                 │ Network | Monitoring | Evaluation      │
                 └────────────────────────────────────────┘

This is not one fixed architecture.

It's a reference model.


61. Architecture Decision: Centralized vs Distributed AI Platform

An enterprise may have two approaches.

Centralized

All AI applications
        │
        ▼
Central AI Platform

Advantages:

  • common controls

  • model governance

  • centralized observability

  • consistent security

Distributed

Application A → AI stack A
Application B → AI stack B
Application C → AI stack C

Advantages:

  • application-specific optimization

  • team autonomy

Trade-off:

  • duplicated infrastructure

  • inconsistent controls

  • fragmented governance

A common enterprise compromise is:

Central platform capabilities
+
Application-specific logic

62. What Should Be Centralized?

Good candidates for centralization often include:

Model Gateway
Identity
Secrets
Policy
Audit
Observability
Evaluation infrastructure
Common retrieval infrastructure

Application-specific components may include:

Business workflow
Domain prompts
Domain tools
Domain knowledge

This gives teams flexibility without allowing every team to reinvent security and platform controls.


63. What Should Be Stateless?

Whenever practical, keep request-serving components stateless:

API
Agent workers
RAG service

Persist state externally:

Database
Cache
Queue
Object storage

This simplifies horizontal scaling.

For example:

                  Load Balancer
                       │
             ┌─────────┼─────────┐
             ▼         ▼         ▼
          Worker A  Worker B  Worker C
             │         │         │
             └─────────┼─────────┘
                       ▼
                  Shared State

64. What Should Be Durable?

Not everything should live only in memory.

Durable information might include:

User / tenant configuration
Agent state
Task state
Audit records
Source documents
Evaluation results
Business transactions

This is especially important when an agent executes a task over several minutes or hours.


65. Long-Running Agent Architecture

For complex tasks:

User
 ↓
Create Job
 ↓
Queue
 ↓
Agent Worker
 ↓
Checkpoint State
 ↓
Tool Calls
 ↓
Checkpoint State
 ↓
More Tool Calls
 ↓
Complete

If Worker A crashes:

Worker B
 ↓
Resume from checkpoint

This is much more robust than storing everything in a process's memory.


66. Security Architecture: The Three Most Important Boundaries

In AI systems, pay special attention to these boundaries.

Boundary 1 — User → AI

Controls:

Authentication
Input validation
Rate limiting

Boundary 2 — AI → Data

Controls:

Authorization
Tenant isolation
DLP
Retrieval filtering

Boundary 3 — AI → Action

Controls:

Tool authorization
Parameter validation
Policy
Human approval
Audit

Visualized:

User
 │
 ▼
[ AI SYSTEM ]
 │
 ├──────────→ Data
 │
 └──────────→ Actions

Each arrow should have explicit controls.


67. Architecture Principle: Least Privilege

This applies everywhere.

Not just users.

Also:

Agents
Tools
Services
Models
Workers
Databases
Caches

For example:

Security Triage Agent
        │
        ├── READ Jira
        ├── READ GitHub
        ├── READ security docs
        └── CREATE Jira comment

        NOT:
        ├── DELETE Jira
        ├── WRITE production DB
        └── READ secrets

68. Architecture Principle: Fail Closed for High-Risk Actions

Suppose an authorization service becomes unavailable.

For a harmless read:

Authorization unavailable
       ↓
Maybe retry

For a destructive production action:

Authorization unavailable
       ↓
DO NOT EXECUTE

This is a classic security architecture principle:

When the cost of an incorrect authorization decision is high, failure should not silently become permission.


69. Architecture Principle: Make AI Decisions Observable

You don't necessarily need to expose internal model reasoning to users.

But you should make system behavior auditable.

For example:

Agent Task: SEC-4321

Tools used:
1. Jira
2. GitHub
3. Scanner

Policy checks:
3 passed
1 blocked

Final action:
Jira comment created

This gives operators visibility without requiring the system to expose private model reasoning.


70. Architecture Interview Question: "Why Did You Choose This Architecture?"

A strong answer should sound like:

"I started from workload characteristics rather than technology preferences. The application requires enterprise knowledge retrieval, occasional multi-step actions, strong tenant isolation, and auditable operations. Therefore I separated deterministic application controls from probabilistic AI reasoning, used a dedicated retrieval layer, introduced a controlled tool gateway for agents, and centralized identity, policy, observability, and cost controls."

That's architecture reasoning.

Not:

"We used Kubernetes because Kubernetes is scalable."


71. Architecture Interview Question: "How Do You Design for Scale?"

A strong answer:

"I first identify the bottleneck rather than scaling every component equally. I separate stateless APIs from durable state, use queues for asynchronous workloads, scale workers horizontally, cache expensive operations where safe, control downstream model and tool quotas, and monitor end-to-end latency and throughput."


72. Architecture Interview Question: "How Do You Control Cost?"

A strong answer:

"I optimize the entire request path: retrieval quality, context size, model routing, caching, agent step limits, batching, and token budgets. I measure cost per request and per workflow rather than looking only at the monthly provider bill."


73. Architecture Interview Question: "How Do You Secure an AI Platform?"

A strong answer:

"I treat the LLM as an untrusted probabilistic component rather than a security boundary. Identity and authorization are enforced outside the model, tools are least-privileged and policy-controlled, tenant isolation is enforced throughout retrieval and storage, sensitive data is controlled across the lifecycle, and agent actions are audited and bounded."


74. Architecture Interview Question: "How Do You Handle a Model Outage?"

A good answer should distinguish between workloads.

For a low-risk chatbot:

Primary model
     ↓
Fallback model
     ↓
Response

For a high-risk security workflow:

Primary model unavailable
        ↓
Fallback unavailable
        ↓
Safe failure
        ↓
Human escalation

The appropriate fallback depends on the risk.


75. Architecture Interview Question: "How Do You Move From POC to Production?"

A strong answer:

"I revisit the design across six dimensions: security, reliability, scale, cost, observability, and governance. I replace local state with durable services where needed, introduce authentication and authorization, productionize ingestion and indexing, add monitoring and evaluation, establish CI/CD and versioning, define failure and recovery behavior, and validate the system under realistic workloads."


76. A Practical Example — Production Security Copilot

Let's imagine we are building:

Enterprise Security Copilot

Capabilities:

  • search security standards

  • inspect Jira vulnerabilities

  • inspect GitHub pull requests

  • run approved security checks

  • summarize evidence

  • create a Jira comment

  • generate a security report

Architecture:

                         Security Engineer
                                  │
                                  ▼
                           Security Copilot
                                  │
                         ┌────────┼─────────┐
                         ▼        ▼         ▼
                        RAG      Agent     Memory
                         │        │
                         │        ▼
                         │    Tool Gateway
                         │        │
                         │   ┌────┼────┐
                         │   ▼    ▼    ▼
                         │ Jira GitHub Scanner
                         │
                         ▼
                    Security KB
                         │
                    Vector + Search

Security controls:

Identity
 ↓
Tenant / RBAC
 ↓
Tool Authorization
 ↓
Policy Engine
 ↓
Approval for sensitive actions
 ↓
Audit

Operational controls:

Tracing
Metrics
Cost controls
Token budgets
Agent step limits
Retries
Timeouts

Now this is an enterprise architecture rather than a chatbot.


77. A Production Readiness Checklist

Before calling an AI application "production ready", ask:

Architecture

□ Clear component boundaries
□ Appropriate data stores
□ Stateless scaling where possible
□ Durable state where required
□ Failure modes defined

AI

□ Model selected using evaluation
□ Prompt/version management
□ RAG evaluation
□ Agent evaluation
□ Model upgrade process

Security

□ Authentication
□ Authorization
□ Tenant isolation
□ Secrets management
□ DLP / sensitive data controls
□ Tool security
□ Prompt injection testing
□ Audit

Operations

□ Metrics
□ Tracing
□ Logging
□ Alerting
□ Cost monitoring
□ Rate limiting
□ Capacity planning

Resilience

□ Timeouts
□ Retry policy
□ Circuit breakers
□ Queueing where needed
□ Backups
□ Disaster recovery

78. The Architecture Principles I Would Remember

There are hundreds of individual design decisions.

But these principles capture most of them.

Principle 1

Don't make the LLM your security boundary.

Principle 2

Separate reasoning, execution, and authorization.

Principle 3

Treat retrieval as a production data pipeline.

Principle 4

Treat prompts, models, indexes, and policies as versioned artifacts.

Principle 5

Give agents narrow capabilities instead of unrestricted power.

Principle 6

Every autonomous workflow needs limits.

Principle 7

Optimize cost and latency across the complete request path.

Principle 8

Monitor quality, not just infrastructure health.

Principle 9

Design for failure before designing for scale.

Principle 10

Use deterministic controls around probabilistic components.


79. The Production AI Mental Model

At this point, the entire series can be summarized as:

                         ENTERPRISE AI

                              USER
                                │
                                ▼
                         APPLICATION
                                │
                    ┌───────────┴───────────┐
                    ▼                       ▼
                   RAG                    AGENT
                    │                       │
               KNOWLEDGE                  ACTION
                    │                       │
                    └───────────┬───────────┘
                                ▼
                               LLM
                                │
                                ▼
                    ┌───────────────────────┐
                    │ Deterministic Controls│
                    │                       │
                    │ IAM                    │
                    │ Authorization         │
                    │ Policy                 │
                    │ Validation             │
                    │ Rate Limits            │
                    │ Cost Controls          │
                    │ Audit                  │
                    └──────────┬────────────┘
                               │
                               ▼
                     ENTERPRISE SYSTEMS

The model provides intelligence.

The surrounding architecture provides:

Control
Security
Reliability
Scale
Governance

That distinction is fundamental.


80. Final Takeaways

A production AI architecture is not:

LLM
+
Vector DB
+
Chat UI

It is a complete distributed system.

You need to reason about:

                 Production AI
                      │
      ┌───────────────┼────────────────┐
      ▼               ▼                ▼
   Intelligence     Data            Execution
      │               │                │
     LLM             RAG              Tools
      │               │                │
      └───────────────┼────────────────┘
                      ▼
                   Platform
                      │
      ┌───────────────┼─────────────────┐
      ▼               ▼                 ▼
   Security       Observability       Reliability
      │               │                 │
      └───────────────┼─────────────────┘
                      ▼
                     Scale
                      │
                      ▼
                     Cost

The most important architectural lesson is:

A production AI system is a distributed application with an AI component—not an AI model surrounded by a few APIs.

And one final principle ties the entire series together:

Use AI where it provides intelligence; use deterministic software where you need guarantees.

Let the model reason.

Let retrieval provide knowledge.

Let agents coordinate actions.

But let identity, authorization, policies, validation, transactions, and security controls remain deterministic and enforceable outside the model.


What's Next?

We've now covered the architecture of:

LLMs
   ↓
RAG
   ↓
Agents
   ↓
Production AI Platform

The obvious next question is the one that matters most to security engineers:

"What can actually go wrong?"

A production AI system introduces a new set of attack surfaces:

User
 ↓
Prompt
 ↓
LLM
 ↓
RAG
 ↓
Memory
 ↓
Tools
 ↓
Enterprise Systems

An attacker may try to manipulate:

Prompts
Documents
Embeddings
Memory
Tool calls
Model outputs

And an innocent failure can be just as dangerous as an intentional attack when an autonomous agent has access to real systems.

That brings us to Part 5 — AI Security: Threat Modeling and Securing LLM, RAG & Agentic AI Applications, where we'll take the architecture above and attack it from a security-engineering perspective.

We'll cover:

  • Prompt injection and indirect prompt injection

  • RAG poisoning

  • Sensitive information disclosure

  • Cross-tenant data leakage

  • Excessive agency

  • Agent/tool abuse

  • Output injection

  • Memory poisoning

  • AI supply-chain security

  • Model and plugin trust

  • SSRF and network risks in AI agents

  • Authorization architecture

  • Human-in-the-loop security

  • AI red teaming

  • Security test cases

  • AI threat modeling

  • Secure reference architecture

  • A complete enterprise AI security checklist