AI Engineering Interview Guide — Part 1: LLM Fundamentals
Tokens, Tokenization, Embeddings, Vectors, Context Windows, Temperature, Top-K, Top-P, and Fine-Tuning — explained with practical examples
Artificial intelligence interviews have changed dramatically.
A few years ago, an AI/ML interview might focus heavily on traditional machine-learning concepts, algorithms, and statistics. Today, for an AI Engineer, GenAI Engineer, AI Security Engineer, or AI Architect, you are increasingly expected to understand how modern LLM applications actually work.
You may be asked questions such as:
What exactly is a token?
What is an embedding?
What is a vector?
How is tokenization different from chunking?
Why do we need embeddings for RAG?
What is a context window?
What does temperature actually do?
What is the difference between Top-K and Top-P?
When should you use RAG instead of fine-tuning?
Can an AI application work without a vector database?
These questions look simple, but they become much easier once you understand the underlying architecture.
This article builds that foundation.
1. What Exactly Is an LLM?
Let's start with the most basic question.
What is an LLM?
An LLM (Large Language Model) is a model trained on large amounts of text so that it can understand and generate language.
At a simplified level:
User Prompt
│
▼
┌──────────────┐
│ LLM │
└──────┬───────┘
│
▼
Generated Output
For example:
User:
"Explain what SQL injection is."
↓
LLM
↓
"SQL injection is a vulnerability where..."
But an important point is often misunderstood:
An LLM does not simply "look up" an answer from a database every time you ask a question.
The model generates output based on patterns and representations learned during training, combined with the information provided in the current context.
This distinction becomes extremely important when we later discuss RAG.
2. What Does an LLM Actually Process?
An LLM doesn't directly process text as humans see it.
It processes tokens.
Consider:
"Application security is important."
The tokenizer converts that text into tokens.
Conceptually:
Application | security | is | important | .
Those tokens are then mapped to numerical IDs.
For example:
Text
│
▼
Tokenizer
│
▼
Token IDs
│
▼
LLM
The exact token boundaries depend on the tokenizer and model.
A token might be:
a complete word
part of a word
punctuation
whitespace-related text
a special token
So the statement:
"One word = one token"
is not generally correct.
3. What Is Tokenization?
Tokenization is the process of breaking text into tokens that the model can process.
For example:
"Secure APIs properly"
might be represented conceptually as:
["Secure", " APIs", " properly"]
The tokenizer then maps those pieces to token IDs:
["Secure", " APIs", " properly"]
│
▼
[1542, 8271, 9132]
The exact numbers are model-specific.
The LLM works with those numerical representations rather than directly processing the original characters.
4. Why Do Tokens Matter?
Tokens are important for three major reasons:
1. Context window
LLMs have a maximum amount of context they can process at once.
2. Cost
Many commercial LLM APIs charge based partly on input and output tokens.
3. Performance
Longer prompts mean more data that the system has to process.
For example:
Prompt A
500 tokens
Prompt B
20,000 tokens
Prompt B may consume significantly more resources and can introduce unnecessary context.
This is why AI engineers care deeply about:
Token efficiency
↓
Cost
↓
Latency
↓
Scalability
5. What Is the Difference Between Tokens and Words?
This is a common interview question.
A word is a human linguistic concept.
A token is a unit defined by the tokenizer.
For example:
"unbelievable"
might be represented as:
"un" + "believ" + "able"
depending on the tokenizer.
Therefore:
Token count and word count are not the same thing.
This matters when calculating:
model cost
prompt size
context limits
output limits
6. What Is an Embedding?
Now we move to one of the most important concepts in modern AI applications.
An embedding is a numerical representation of information such as text that attempts to capture its semantic characteristics.
For example:
"How do I reset my password?"
is converted by an embedding model into something like:
[0.12, -0.38, 0.77, 0.24, ...]
This is called an embedding vector.
The vector may contain hundreds or thousands of dimensions depending on the embedding model.
The important idea is not the individual numbers.
The important idea is:
Similar meanings should generally produce embeddings that are close to one another according to the chosen similarity metric.
7. Practical Example of Embeddings
Imagine we have these three sentences:
A = "How can I reset my password?"
B = "I forgot my password. How do I change it?"
C = "What is the weather in Mumbai today?"
The embeddings might conceptually look like:
A → [0.11, 0.73, 0.21, ...]
B → [0.13, 0.71, 0.23, ...]
C → [-0.82, 0.14, 0.66, ...]
The distance between A and B is likely much smaller than between A and C.
Password Questions
● A
● B
● C
Weather
That ability to compare semantic similarity is what makes embeddings so useful for:
semantic search
RAG
document retrieval
recommendation systems
clustering
duplicate detection
8. What Is a Vector?
A vector is simply an ordered collection of numerical values.
For example:
[0.12, -0.38, 0.77, 0.24]
In AI systems, vectors are commonly used to represent things such as:
text
images
audio
documents
users
products
In a RAG system, an embedding is usually represented as a vector.
So you can think of:
Text
↓
Embedding model
↓
Embedding
↓
Vector
In casual discussions, people sometimes use embedding and vector almost interchangeably.
Technically, though:
Embedding describes the learned representation, while vector describes the numerical form used to represent it.
9. How Does Semantic Search Work?
Suppose your company has 100,000 security documents.
A user asks:
"What MFA controls are required for administrator accounts?"
Searching for exact words may miss documents that use different wording.
For example, a document might say:
"Privileged identities must use phishing-resistant multi-factor authentication."
There is not an exact keyword match for the user's sentence, but the concepts are closely related.
Embeddings help solve this.
User Query
│
▼
Query Embedding
│
▼
Similarity Search
│
┌───────────┼───────────┐
▼ ▼ ▼
Doc A Doc B Doc C
0.93 0.88 0.42
The system can retrieve the semantically relevant documents.
This is one of the foundations of RAG.
10. What Is Chunking?
Now we need another concept that is often confused with tokenization.
Suppose you have a 500-page security policy.
You usually don't want to treat the entire document as one retrieval unit.
Instead, you divide it into smaller pieces called chunks.
500-page document
│
▼
Chunking
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Chunk 1 Chunk 2 Chunk 3
For example:
Chunk 1
Authentication requirements
Chunk 2
MFA requirements
Chunk 3
Password requirements
Chunk 4
Session management
These chunks can then be embedded and indexed for retrieval.
11. Tokenization vs Chunking
This is one of the easiest interview questions to get wrong.
Tokenization
Breaks text into units that the model can process.
Text
↓
Tokens
Chunking
Breaks a document into meaningful retrieval units.
Document
↓
Chunks
The relationship is:
Document
│
▼
Tokenizer
│
▼
Tokens
│
▼
Chunking Strategy
│
├── Chunk 1
├── Chunk 2
└── Chunk 3
So remember:
Tokenization is model representation. Chunking is information organization for retrieval.
We will go much deeper into chunking strategies in Part 2 of this series.
12. What Is a Context Window?
An LLM can only process a certain amount of information at a time.
This available space is commonly referred to as the context window.
Conceptually:
┌────────────────────────────────────────┐
│ Context Window │
│ │
│ System Instructions │
│ + │
│ User Prompt │
│ + │
│ Conversation History │
│ + │
│ Retrieved Documents │
│ │
└────────────────────────────────────────┘
For example, a system may need to provide:
System instructions
+
User question
+
5 retrieved documents
+
Conversation history
All of this consumes context.
13. Why Is Context Important?
Imagine asking an AI security assistant:
"Review this security architecture."
You give it:
20 pages of architecture
+
30 pages of logs
+
10 pages of requirements
+
15 pages of historical tickets
Even if the model supports a large context window, throwing everything into the prompt is not automatically a good architecture.
More context can result in:
higher cost
increased latency
irrelevant information
conflicting information
context dilution
more difficult reasoning
This leads to a very important engineering principle:
The goal is not to provide the maximum amount of context. The goal is to provide the right context.
That principle becomes central to RAG design.
14. What Is Temperature?
Temperature controls the randomness of token selection during generation.
Conceptually:
Low Temperature
↓
More predictable output
High Temperature
↓
More variation
Imagine asking:
"What is SQL injection?"
With a lower temperature:
SQL injection is a vulnerability where...
The output tends to be more consistent.
With a higher temperature, the model may produce more stylistic variation.
15. Does Higher Temperature Make the Model Smarter?
No.
This is a common misconception.
Temperature controls sampling behavior, not intelligence.
Think of it as:
Temperature
↓
How adventurous is the generation?
not:
Temperature
↓
How intelligent is the model?
For example:
Security classification system
You may want:
Low temperature
because predictable output is desirable.
Creative writing application
You may intentionally choose:
Higher temperature
because variation is desirable.
16. What Is Top-K Sampling?
Top-K is a generation parameter.
Suppose the model predicts:
The vulnerability is caused by ...
SQL 35%
authentication 25%
improper 15%
weak 10%
...
With:
Top-K = 3
the system considers only the top three candidate tokens.
SQL
authentication
improper
Then sampling happens among those candidates.
The critical point:
Generation Top-K is about selecting candidate output tokens.
It is completely different from Top-K document retrieval in RAG.
17. What Is Top-P Sampling?
Top-P is also a generation parameter, but it works differently.
Instead of saying:
"Keep exactly 10 tokens."
you say:
"Keep the smallest group of tokens whose cumulative probability reaches P."
For example:
Token probabilities:
A = 45%
B = 25%
C = 15%
D = 8%
E = 4%
F = 3%
With:
Top-P = 0.85
the system keeps enough candidates to reach approximately 85% cumulative probability.
So:
45 + 25 + 15 = 85%
and the remaining candidates are excluded.
18. Top-K vs Top-P
This is a classic interview comparison.
| Parameter | Meaning |
|---|---|
| Temperature | Controls randomness |
| Top-K | Keep the K most probable token candidates |
| Top-P | Keep candidates up to a cumulative probability threshold |
And remember:
RAG Top-K
≠
Generation Top-K
RAG retrieval
≠
Token sampling
For example:
RAG:
Query
↓
Retrieve Top-K documents
↓
LLM
Generation:
LLM
↓
Top-K / Top-P candidate tokens
↓
Output
This distinction is worth remembering.
19. A Practical Example: Building a Security Assistant
Suppose we are building an internal AI security assistant.
The user asks:
"Does our API gateway require MFA for privileged access?"
A simplistic application might simply send the question to an LLM:
User
↓
LLM
↓
Answer
But the LLM may not know the company's latest internal policy.
Instead, we can build:
User
│
▼
User Query
│
▼
Retrieval System
│
▼
Relevant Documents
│
▼
LLM
│
▼
Answer
The documents might contain:
Security Policy v2026
Identity Standard
Privileged Access Standard
API Security Standard
This is the basic idea behind Retrieval-Augmented Generation (RAG).
We will explore this architecture in detail in Part 2.
20. Can We Work With AI Without a Vector Database?
Yes.
This is another very useful interview question.
A vector database is not mandatory for every AI application.
You can use:
Traditional keyword search
For example:
BM25
Elasticsearch
OpenSearch
Relational databases
For structured data:
PostgreSQL
MySQL
SQL Server
Graph databases
For highly connected relationships:
User
↓ owns
Account
↓ contains
Portfolio
Hybrid retrieval
You can combine:
Keyword Search
+
Semantic Search
↓
Reranking
A mature RAG architecture may therefore look like:
Query
│
┌────────┴────────┐
▼ ▼
Keyword Search Vector Search
│ │
└────────┬────────┘
▼
Reranker
│
▼
Relevant Context
The correct database depends on the problem.
21. Why Do We Need Embeddings If We Already Have an LLM?
This is a subtle but important question.
An LLM and an embedding model serve different purposes.
LLM
Primarily used for things such as:
understanding prompts
reasoning
generation
summarization
transformation
Embedding model
Primarily used to create representations suitable for similarity search.
For example:
"Reset password"
│
▼
Embedding Model
│
▼
Vector [....]
│
▼
Similarity Search
You could use the same vendor for both, but conceptually they are different components.
22. What Is Fine-Tuning?
Fine-tuning means taking a pre-trained model and further training it on a task-specific dataset.
Conceptually:
Pre-trained Model
│
▼
Task-specific Dataset
│
▼
Fine-tuning
│
▼
Specialized Model
For example, imagine you want a model to consistently classify application-security findings into:
SQL Injection
XSS
SSRF
Authentication
Authorization
Cryptography
You could fine-tune a model using examples of correctly classified findings.
23. Fine-Tuning vs RAG
This is one of the most important architecture questions.
Suppose your company has:
Security Policy 2025
Security Policy 2026
Security Policy 2027
and the model needs to answer questions about the latest policy.
Would you immediately fine-tune the model?
Usually, no.
The problem here is primarily:
The model needs access to changing information.
RAG is often better suited:
Question
↓
Retrieve latest policy
↓
LLM
↓
Answer
Now consider another problem:
"We want the model to consistently output vulnerability reports in our specific internal format."
That may be a better candidate for fine-tuning, depending on the use case.
So a useful decision framework is:
Need better AI behavior?
│
▼
Try prompting
│
▼
Need external/current knowledge?
│
▼
RAG
│
▼
Need consistent specialized behavior?
│
▼
Consider fine-tuning
The important lesson is:
RAG gives the model access to information. Fine-tuning changes the model's learned behavior.
24. A Practical Security Example: RAG vs Fine-Tuning
Imagine you're building an AI assistant for penetration testers.
Requirement A
"Answer using our latest internal secure coding standard."
This information changes regularly.
RAG is a natural fit:
Latest Security Standard
↓
Chunk
↓
Embedding
↓
Retrieval
↓
LLM
Requirement B
"Generate vulnerability reports using our standardized internal reporting style."
This is more about consistent behavior and output format.
Prompting, structured outputs, or potentially fine-tuning may be considered.
The architecture decision should therefore be based on the problem you are solving, not on the assumption that fine-tuning is always "more advanced."
25. What Happens When You Combine Everything?
Now let's connect the concepts.
Imagine an enterprise AI assistant:
User
│
▼
User Prompt
│
▼
Tokens
│
▼
LLM
│
┌──────┴──────┐
│ │
▼ ▼
Context Tools
│
▼
RAG
│
┌────┴─────┐
▼ ▼
Embeddings Search
│
▼
Vector DB
│
▼
Chunks
│
▼
Context
│
▼
LLM
│
▼
Response
This looks complicated initially, but each component has a very specific responsibility.
26. The Mental Model You Should Remember
You don't need to memorize dozens of disconnected definitions.
Build this mental model instead:
DOCUMENT
│
▼
TOKENIZER
│
▼
TOKENS
│
▼
CHUNKING
│
▼
EMBEDDING
│
▼
VECTOR
│
▼
SEARCH / RAG
│
▼
RETRIEVED
CONTEXT
│
▼
LLM
│
┌─────────┴─────────┐
▼ ▼
PARAMETERS OUTPUT
│
Temperature
Top-K
Top-P
Once this makes sense, the next concepts—RAG, agents, tool calling, memory, and AI security—become much easier.
27. Interview Cheat Sheet
Here are the concepts in one place.
Token
A unit of text processed by a model.
Tokenization
Converting text into tokens/token IDs.
Chunk
A piece of a larger document used as a retrieval unit.
Embedding
A learned numerical representation capturing useful semantic characteristics.
Vector
The numerical representation used to represent an embedding.
Context Window
The amount of information the model can process in a single context.
Temperature
Controls randomness/variation during generation.
Top-K
Limits generation to the K highest-probability token candidates.
Top-P
Limits generation to the smallest probability set whose cumulative probability reaches P.
RAG
Retrieves external/contextual information and provides it to the LLM before generation.
Fine-Tuning
Further trains a model to specialize its behavior for a particular task/domain.
28. Quick Interview Questions
Before moving to RAG, make sure you can answer these without memorizing a script.
Q1. What is a token?
A token is a unit of text processed by the model. It may represent a word, part of a word, punctuation, or another tokenization unit.
Q2. What is an embedding?
An embedding is a numerical representation of data, such as text, designed to capture useful semantic relationships.
Q3. What is a vector?
A vector is an ordered set of numerical values. Embeddings are commonly stored and searched as vectors.
Q4. Tokenization vs chunking?
Tokenization converts text into model-readable tokens. Chunking divides documents into retrieval-sized pieces.
Q5. Why do we need embeddings?
They enable semantic similarity and retrieval, such as finding documents that mean something similar even when they don't contain the exact same words.
Q6. Does RAG require a vector database?
No. RAG can use vector search, keyword search, hybrid search, databases, graph stores, or other retrieval mechanisms.
Q7. Does Top-K always mean RAG?
No. Top-K can refer to document retrieval in RAG or token sampling during LLM generation. Context matters.
Q8. Does higher temperature mean a smarter model?
No. Temperature mainly changes the randomness/variation of generation.
Q9. RAG vs fine-tuning?
RAG primarily provides external/current knowledge at inference time. Fine-tuning changes or specializes model behavior through additional training.
Q10. What should an AI architect optimize for?
Not simply "the biggest model."
A production system must consider:
Quality
+
Security
+
Latency
+
Cost
+
Reliability
+
Scalability
+
Data Privacy
29. The Most Important Takeaways
If you remember only a few things from this article, remember these:
Tokens are how the model processes text.
Chunking is how applications divide information for retrieval.
Embeddings turn information into numerical representations that can be compared semantically.
Vectors are the numerical form used for those representations.
Context is what the model can see for the current task.
Temperature, Top-K, and Top-P influence how the model generates output.
RAG retrieves information; it does not retrain the model.
Fine-tuning changes model behavior; it is not simply another way to store documents.
And perhaps the most useful engineering principle:
Don't choose an AI technology because it sounds advanced. Choose the architecture that solves the actual problem.
What's Next?
In this article, we established the building blocks:
Tokens
↓
Embeddings
↓
Vectors
↓
Chunks
↓
Retrieval
↓
LLM
The next question naturally becomes:
How do we take thousands or millions of enterprise documents, break them into useful chunks, convert them into embeddings, retrieve the right information, and give only the relevant context to the LLM?
That's where RAG comes in.
In Part 2 — RAG Deep Dive, we'll go beyond the definition of RAG and walk through the complete production pipeline:
Documents
↓
Parsing
↓
Chunking
↓
Embedding
↓
Vector / Keyword Index
↓
Query
↓
Retrieval
↓
Top-K
↓
Reranking
↓
Context Construction
↓
LLM
↓
Answer
We'll also tackle the questions AI engineers are frequently asked in interviews:
How do you choose chunk size?
Fixed-size vs semantic chunking?
How much overlap should you use?
What is hybrid retrieval?
What is reranking?
How do you measure retrieval quality?
Why can retrieving too much information make RAG worse?
Can you build RAG without a vector database?
How do you deal with stale documents?
How do you prevent one tenant from retrieving another tenant's data?
And eventually, we'll take the next step:
From an AI that retrieves information to an AI agent that can decide what to do and take actions.

