Skip to main content

Command Palette

Search for a command to run...

AI Engineering Interview Guide — Part 1: LLM Fundamentals

Updated
•19 min read•View as Markdown
S
I like breaking things that are supposed to be secure. When I’m not hunting vulnerabilities, I’m exploring systems, architectures, and the assumptions behind them.

Tokens, Tokenization, Embeddings, Vectors, Context Windows, Temperature, Top-K, Top-P, and Fine-Tuning — explained with practical examples

Artificial intelligence interviews have changed dramatically.

A few years ago, an AI/ML interview might focus heavily on traditional machine-learning concepts, algorithms, and statistics. Today, for an AI Engineer, GenAI Engineer, AI Security Engineer, or AI Architect, you are increasingly expected to understand how modern LLM applications actually work.

You may be asked questions such as:

  • What exactly is a token?

  • What is an embedding?

  • What is a vector?

  • How is tokenization different from chunking?

  • Why do we need embeddings for RAG?

  • What is a context window?

  • What does temperature actually do?

  • What is the difference between Top-K and Top-P?

  • When should you use RAG instead of fine-tuning?

  • Can an AI application work without a vector database?

These questions look simple, but they become much easier once you understand the underlying architecture.

This article builds that foundation.


1. What Exactly Is an LLM?

Let's start with the most basic question.

What is an LLM?

An LLM (Large Language Model) is a model trained on large amounts of text so that it can understand and generate language.

At a simplified level:

                User Prompt
                     │
                     ▼
             ┌──────────────┐
             │     LLM      │
             └──────┬───────┘
                    │
                    ▼
             Generated Output

For example:

User:
"Explain what SQL injection is."

             ↓

LLM

             ↓

"SQL injection is a vulnerability where..."

But an important point is often misunderstood:

An LLM does not simply "look up" an answer from a database every time you ask a question.

The model generates output based on patterns and representations learned during training, combined with the information provided in the current context.

This distinction becomes extremely important when we later discuss RAG.


2. What Does an LLM Actually Process?

An LLM doesn't directly process text as humans see it.

It processes tokens.

Consider:

"Application security is important."

The tokenizer converts that text into tokens.

Conceptually:

Application | security | is | important | .

Those tokens are then mapped to numerical IDs.

For example:

Text
 │
 ▼
Tokenizer
 │
 ▼
Token IDs
 │
 ▼
LLM

The exact token boundaries depend on the tokenizer and model.

A token might be:

  • a complete word

  • part of a word

  • punctuation

  • whitespace-related text

  • a special token

So the statement:

"One word = one token"

is not generally correct.


3. What Is Tokenization?

Tokenization is the process of breaking text into tokens that the model can process.

For example:

"Secure APIs properly"

might be represented conceptually as:

["Secure", " APIs", " properly"]

The tokenizer then maps those pieces to token IDs:

["Secure", " APIs", " properly"]
          │
          ▼
      [1542, 8271, 9132]

The exact numbers are model-specific.

The LLM works with those numerical representations rather than directly processing the original characters.


4. Why Do Tokens Matter?

Tokens are important for three major reasons:

1. Context window

LLMs have a maximum amount of context they can process at once.

2. Cost

Many commercial LLM APIs charge based partly on input and output tokens.

3. Performance

Longer prompts mean more data that the system has to process.

For example:

Prompt A
500 tokens

Prompt B
20,000 tokens

Prompt B may consume significantly more resources and can introduce unnecessary context.

This is why AI engineers care deeply about:

Token efficiency
     ↓
Cost
     ↓
Latency
     ↓
Scalability

5. What Is the Difference Between Tokens and Words?

This is a common interview question.

A word is a human linguistic concept.

A token is a unit defined by the tokenizer.

For example:

"unbelievable"

might be represented as:

"un" + "believ" + "able"

depending on the tokenizer.

Therefore:

Token count and word count are not the same thing.

This matters when calculating:

  • model cost

  • prompt size

  • context limits

  • output limits


6. What Is an Embedding?

Now we move to one of the most important concepts in modern AI applications.

An embedding is a numerical representation of information such as text that attempts to capture its semantic characteristics.

For example:

"How do I reset my password?"

is converted by an embedding model into something like:

[0.12, -0.38, 0.77, 0.24, ...]

This is called an embedding vector.

The vector may contain hundreds or thousands of dimensions depending on the embedding model.

The important idea is not the individual numbers.

The important idea is:

Similar meanings should generally produce embeddings that are close to one another according to the chosen similarity metric.


7. Practical Example of Embeddings

Imagine we have these three sentences:

A = "How can I reset my password?"

B = "I forgot my password. How do I change it?"

C = "What is the weather in Mumbai today?"

The embeddings might conceptually look like:

A → [0.11, 0.73, 0.21, ...]
B → [0.13, 0.71, 0.23, ...]
C → [-0.82, 0.14, 0.66, ...]

The distance between A and B is likely much smaller than between A and C.

          Password Questions
                 ● A
                ● B


                          ● C
                       Weather

That ability to compare semantic similarity is what makes embeddings so useful for:

  • semantic search

  • RAG

  • document retrieval

  • recommendation systems

  • clustering

  • duplicate detection


8. What Is a Vector?

A vector is simply an ordered collection of numerical values.

For example:

[0.12, -0.38, 0.77, 0.24]

In AI systems, vectors are commonly used to represent things such as:

  • text

  • images

  • audio

  • documents

  • users

  • products

In a RAG system, an embedding is usually represented as a vector.

So you can think of:

Text
  ↓
Embedding model
  ↓
Embedding
  ↓
Vector

In casual discussions, people sometimes use embedding and vector almost interchangeably.

Technically, though:

Embedding describes the learned representation, while vector describes the numerical form used to represent it.


9. How Does Semantic Search Work?

Suppose your company has 100,000 security documents.

A user asks:

"What MFA controls are required for administrator accounts?"

Searching for exact words may miss documents that use different wording.

For example, a document might say:

"Privileged identities must use phishing-resistant multi-factor authentication."

There is not an exact keyword match for the user's sentence, but the concepts are closely related.

Embeddings help solve this.

                  User Query
                      │
                      ▼
               Query Embedding
                      │
                      ▼
               Similarity Search
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
       Doc A        Doc B        Doc C
       0.93         0.88         0.42

The system can retrieve the semantically relevant documents.

This is one of the foundations of RAG.


10. What Is Chunking?

Now we need another concept that is often confused with tokenization.

Suppose you have a 500-page security policy.

You usually don't want to treat the entire document as one retrieval unit.

Instead, you divide it into smaller pieces called chunks.

               500-page document
                       │
                       ▼
                  Chunking
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
     Chunk 1        Chunk 2        Chunk 3

For example:

Chunk 1
Authentication requirements

Chunk 2
MFA requirements

Chunk 3
Password requirements

Chunk 4
Session management

These chunks can then be embedded and indexed for retrieval.


11. Tokenization vs Chunking

This is one of the easiest interview questions to get wrong.

Tokenization

Breaks text into units that the model can process.

Text
 ↓
Tokens

Chunking

Breaks a document into meaningful retrieval units.

Document
 ↓
Chunks

The relationship is:

Document
   │
   ▼
Tokenizer
   │
   ▼
Tokens
   │
   ▼
Chunking Strategy
   │
   ├── Chunk 1
   ├── Chunk 2
   └── Chunk 3

So remember:

Tokenization is model representation. Chunking is information organization for retrieval.

We will go much deeper into chunking strategies in Part 2 of this series.


12. What Is a Context Window?

An LLM can only process a certain amount of information at a time.

This available space is commonly referred to as the context window.

Conceptually:

┌────────────────────────────────────────┐
│             Context Window              │
│                                        │
│ System Instructions                    │
│ +                                      │
│ User Prompt                            │
│ +                                      │
│ Conversation History                   │
│ +                                      │
│ Retrieved Documents                    │
│                                        │
└────────────────────────────────────────┘

For example, a system may need to provide:

System instructions
+
User question
+
5 retrieved documents
+
Conversation history

All of this consumes context.


13. Why Is Context Important?

Imagine asking an AI security assistant:

"Review this security architecture."

You give it:

20 pages of architecture
+
30 pages of logs
+
10 pages of requirements
+
15 pages of historical tickets

Even if the model supports a large context window, throwing everything into the prompt is not automatically a good architecture.

More context can result in:

  • higher cost

  • increased latency

  • irrelevant information

  • conflicting information

  • context dilution

  • more difficult reasoning

This leads to a very important engineering principle:

The goal is not to provide the maximum amount of context. The goal is to provide the right context.

That principle becomes central to RAG design.


14. What Is Temperature?

Temperature controls the randomness of token selection during generation.

Conceptually:

Low Temperature
       ↓
More predictable output

High Temperature
       ↓
More variation

Imagine asking:

"What is SQL injection?"

With a lower temperature:

SQL injection is a vulnerability where...

The output tends to be more consistent.

With a higher temperature, the model may produce more stylistic variation.


15. Does Higher Temperature Make the Model Smarter?

No.

This is a common misconception.

Temperature controls sampling behavior, not intelligence.

Think of it as:

Temperature
     ↓
How adventurous is the generation?

not:

Temperature
     ↓
How intelligent is the model?

For example:

Security classification system

You may want:

Low temperature

because predictable output is desirable.

Creative writing application

You may intentionally choose:

Higher temperature

because variation is desirable.


16. What Is Top-K Sampling?

Top-K is a generation parameter.

Suppose the model predicts:

The vulnerability is caused by ...

SQL          35%
authentication 25%
improper     15%
weak         10%
...

With:

Top-K = 3

the system considers only the top three candidate tokens.

SQL
authentication
improper

Then sampling happens among those candidates.

The critical point:

Generation Top-K is about selecting candidate output tokens.

It is completely different from Top-K document retrieval in RAG.


17. What Is Top-P Sampling?

Top-P is also a generation parameter, but it works differently.

Instead of saying:

"Keep exactly 10 tokens."

you say:

"Keep the smallest group of tokens whose cumulative probability reaches P."

For example:

Token probabilities:

A = 45%
B = 25%
C = 15%
D = 8%
E = 4%
F = 3%

With:

Top-P = 0.85

the system keeps enough candidates to reach approximately 85% cumulative probability.

So:

45 + 25 + 15 = 85%

and the remaining candidates are excluded.


18. Top-K vs Top-P

This is a classic interview comparison.

Parameter Meaning
Temperature Controls randomness
Top-K Keep the K most probable token candidates
Top-P Keep candidates up to a cumulative probability threshold

And remember:

RAG Top-K
    ≠
Generation Top-K

RAG retrieval
    ≠
Token sampling

For example:

RAG:

Query
 ↓
Retrieve Top-K documents
 ↓
LLM


Generation:

LLM
 ↓
Top-K / Top-P candidate tokens
 ↓
Output

This distinction is worth remembering.


19. A Practical Example: Building a Security Assistant

Suppose we are building an internal AI security assistant.

The user asks:

"Does our API gateway require MFA for privileged access?"

A simplistic application might simply send the question to an LLM:

User
 ↓
LLM
 ↓
Answer

But the LLM may not know the company's latest internal policy.

Instead, we can build:

                    User
                      │
                      ▼
                 User Query
                      │
                      ▼
               Retrieval System
                      │
                      ▼
              Relevant Documents
                      │
                      ▼
                     LLM
                      │
                      ▼
                   Answer

The documents might contain:

Security Policy v2026
Identity Standard
Privileged Access Standard
API Security Standard

This is the basic idea behind Retrieval-Augmented Generation (RAG).

We will explore this architecture in detail in Part 2.


20. Can We Work With AI Without a Vector Database?

Yes.

This is another very useful interview question.

A vector database is not mandatory for every AI application.

You can use:

For example:

BM25
Elasticsearch
OpenSearch

Relational databases

For structured data:

PostgreSQL
MySQL
SQL Server

Graph databases

For highly connected relationships:

User
 ↓ owns
Account
 ↓ contains
Portfolio

Hybrid retrieval

You can combine:

Keyword Search
       +
Semantic Search
       ↓
    Reranking

A mature RAG architecture may therefore look like:

                 Query
                   │
          ┌────────┴────────┐
          ▼                 ▼
    Keyword Search     Vector Search
          │                 │
          └────────┬────────┘
                   ▼
                Reranker
                   │
                   ▼
             Relevant Context

The correct database depends on the problem.


21. Why Do We Need Embeddings If We Already Have an LLM?

This is a subtle but important question.

An LLM and an embedding model serve different purposes.

LLM

Primarily used for things such as:

  • understanding prompts

  • reasoning

  • generation

  • summarization

  • transformation

Embedding model

Primarily used to create representations suitable for similarity search.

For example:

               "Reset password"
                       │
                       ▼
                Embedding Model
                       │
                       ▼
                  Vector [....]
                       │
                       ▼
                Similarity Search

You could use the same vendor for both, but conceptually they are different components.


22. What Is Fine-Tuning?

Fine-tuning means taking a pre-trained model and further training it on a task-specific dataset.

Conceptually:

Pre-trained Model
       │
       ▼
Task-specific Dataset
       │
       ▼
Fine-tuning
       │
       ▼
Specialized Model

For example, imagine you want a model to consistently classify application-security findings into:

SQL Injection
XSS
SSRF
Authentication
Authorization
Cryptography

You could fine-tune a model using examples of correctly classified findings.


23. Fine-Tuning vs RAG

This is one of the most important architecture questions.

Suppose your company has:

Security Policy 2025
Security Policy 2026
Security Policy 2027

and the model needs to answer questions about the latest policy.

Would you immediately fine-tune the model?

Usually, no.

The problem here is primarily:

The model needs access to changing information.

RAG is often better suited:

Question
 ↓
Retrieve latest policy
 ↓
LLM
 ↓
Answer

Now consider another problem:

"We want the model to consistently output vulnerability reports in our specific internal format."

That may be a better candidate for fine-tuning, depending on the use case.

So a useful decision framework is:

Need better AI behavior?
          │
          ▼
     Try prompting
          │
          ▼
Need external/current knowledge?
          │
          ▼
        RAG
          │
          ▼
Need consistent specialized behavior?
          │
          ▼
   Consider fine-tuning

The important lesson is:

RAG gives the model access to information. Fine-tuning changes the model's learned behavior.


24. A Practical Security Example: RAG vs Fine-Tuning

Imagine you're building an AI assistant for penetration testers.

Requirement A

"Answer using our latest internal secure coding standard."

This information changes regularly.

RAG is a natural fit:

Latest Security Standard
        ↓
      Chunk
        ↓
     Embedding
        ↓
    Retrieval
        ↓
       LLM

Requirement B

"Generate vulnerability reports using our standardized internal reporting style."

This is more about consistent behavior and output format.

Prompting, structured outputs, or potentially fine-tuning may be considered.

The architecture decision should therefore be based on the problem you are solving, not on the assumption that fine-tuning is always "more advanced."


25. What Happens When You Combine Everything?

Now let's connect the concepts.

Imagine an enterprise AI assistant:

                         User
                           │
                           ▼
                      User Prompt
                           │
                           ▼
                        Tokens
                           │
                           ▼
                         LLM
                           │
                    ┌──────┴──────┐
                    │             │
                    ▼             ▼
                 Context        Tools
                    │
                    ▼
                   RAG
                    │
               ┌────┴─────┐
               ▼          ▼
           Embeddings   Search
               │
               ▼
           Vector DB
               │
               ▼
             Chunks
               │
               ▼
             Context
               │
               ▼
               LLM
               │
               ▼
            Response

This looks complicated initially, but each component has a very specific responsibility.


26. The Mental Model You Should Remember

You don't need to memorize dozens of disconnected definitions.

Build this mental model instead:

                     DOCUMENT
                        │
                        ▼
                    TOKENIZER
                        │
                        ▼
                      TOKENS
                        │
                        ▼
                     CHUNKING
                        │
                        ▼
                    EMBEDDING
                        │
                        ▼
                      VECTOR
                        │
                        ▼
                  SEARCH / RAG
                        │
                        ▼
                   RETRIEVED
                    CONTEXT
                        │
                        ▼
                       LLM
                        │
              ┌─────────┴─────────┐
              ▼                   ▼
          PARAMETERS           OUTPUT
              │
       Temperature
       Top-K
       Top-P

Once this makes sense, the next concepts—RAG, agents, tool calling, memory, and AI security—become much easier.


27. Interview Cheat Sheet

Here are the concepts in one place.

Token

A unit of text processed by a model.

Tokenization

Converting text into tokens/token IDs.

Chunk

A piece of a larger document used as a retrieval unit.

Embedding

A learned numerical representation capturing useful semantic characteristics.

Vector

The numerical representation used to represent an embedding.

Context Window

The amount of information the model can process in a single context.

Temperature

Controls randomness/variation during generation.

Top-K

Limits generation to the K highest-probability token candidates.

Top-P

Limits generation to the smallest probability set whose cumulative probability reaches P.

RAG

Retrieves external/contextual information and provides it to the LLM before generation.

Fine-Tuning

Further trains a model to specialize its behavior for a particular task/domain.


28. Quick Interview Questions

Before moving to RAG, make sure you can answer these without memorizing a script.

Q1. What is a token?

A token is a unit of text processed by the model. It may represent a word, part of a word, punctuation, or another tokenization unit.

Q2. What is an embedding?

An embedding is a numerical representation of data, such as text, designed to capture useful semantic relationships.

Q3. What is a vector?

A vector is an ordered set of numerical values. Embeddings are commonly stored and searched as vectors.

Q4. Tokenization vs chunking?

Tokenization converts text into model-readable tokens. Chunking divides documents into retrieval-sized pieces.

Q5. Why do we need embeddings?

They enable semantic similarity and retrieval, such as finding documents that mean something similar even when they don't contain the exact same words.

Q6. Does RAG require a vector database?

No. RAG can use vector search, keyword search, hybrid search, databases, graph stores, or other retrieval mechanisms.

Q7. Does Top-K always mean RAG?

No. Top-K can refer to document retrieval in RAG or token sampling during LLM generation. Context matters.

Q8. Does higher temperature mean a smarter model?

No. Temperature mainly changes the randomness/variation of generation.

Q9. RAG vs fine-tuning?

RAG primarily provides external/current knowledge at inference time. Fine-tuning changes or specializes model behavior through additional training.

Q10. What should an AI architect optimize for?

Not simply "the biggest model."

A production system must consider:

Quality
+
Security
+
Latency
+
Cost
+
Reliability
+
Scalability
+
Data Privacy

29. The Most Important Takeaways

If you remember only a few things from this article, remember these:

Tokens are how the model processes text.

Chunking is how applications divide information for retrieval.

Embeddings turn information into numerical representations that can be compared semantically.

Vectors are the numerical form used for those representations.

Context is what the model can see for the current task.

Temperature, Top-K, and Top-P influence how the model generates output.

RAG retrieves information; it does not retrain the model.

Fine-tuning changes model behavior; it is not simply another way to store documents.

And perhaps the most useful engineering principle:

Don't choose an AI technology because it sounds advanced. Choose the architecture that solves the actual problem.


What's Next?

In this article, we established the building blocks:

Tokens
   ↓
Embeddings
   ↓
Vectors
   ↓
Chunks
   ↓
Retrieval
   ↓
LLM

The next question naturally becomes:

How do we take thousands or millions of enterprise documents, break them into useful chunks, convert them into embeddings, retrieve the right information, and give only the relevant context to the LLM?

That's where RAG comes in.

In Part 2 — RAG Deep Dive, we'll go beyond the definition of RAG and walk through the complete production pipeline:

Documents
   ↓
Parsing
   ↓
Chunking
   ↓
Embedding
   ↓
Vector / Keyword Index
   ↓
Query
   ↓
Retrieval
   ↓
Top-K
   ↓
Reranking
   ↓
Context Construction
   ↓
LLM
   ↓
Answer

We'll also tackle the questions AI engineers are frequently asked in interviews:

  • How do you choose chunk size?

  • Fixed-size vs semantic chunking?

  • How much overlap should you use?

  • What is hybrid retrieval?

  • What is reranking?

  • How do you measure retrieval quality?

  • Why can retrieving too much information make RAG worse?

  • Can you build RAG without a vector database?

  • How do you deal with stale documents?

  • How do you prevent one tenant from retrieving another tenant's data?

And eventually, we'll take the next step:

From an AI that retrieves information to an AI agent that can decide what to do and take actions.