AI

How AI Code Explainers Work: AST Traversal, Scope Resolution, and Embeddings

DD
Ankur Ishwar
7 min read Updated Sep 6, 2026
AST Traversal and Semantic Embeddings for AI Code Explanations

The Problem with Raw Code Prompting

Most developers assume AI code explanation tools work by feeding raw source strings directly into a large language model. You copy an unfamiliar 150-line utility function, send it to a chat prompt, and receive an explanation. When the snippet is trivial, this approach appears to work. On production codebases with complex lexical scopes, asynchronous callbacks, and cross-module dependencies, naive prompting breaks down.

Language models process text as linear sequences of sub-word tokens. Code is fundamentally non-linear: it is a graph of scopes, references, type assertions, and execution branches. When fed raw source text without structural scaffolding, models easily mistake shadowed variables for outer declarations, miss implicit closures, and hallucinate runtime control flows. Production-grade code explainers solve this by pairing Abstract Syntax Tree (AST) traversal with dense contextual embeddings.

The Architecture: From Lexical Tokens to Structural Intelligence

Before any prompt reaches an LLM, the explanation pipeline converts source code into a structured intermediate representation. The AST extracts exact symbol relationships, while vector embeddings locate related definitions across the repository.

Source Code (.ts / .js)
         │
         ▼
[Lexer & Parser (@babel/parser)] ───► Generates Abstract Syntax Tree
         │
         ▼
[AST Visitor (@babel/traverse)]  ───► Extracts: Functions, Scopes, Cyclomatic Complexity
         │
         ▼
[Vector Search (pgvector/LanceDB)] ──► Retrieves related schemas and imported interfaces
         │
         ▼
[LLM Structured Prompt]          ───► Generates line-accurate, hallucination-free explanation

Step 1: Parsing Source Code into an AST

An AST converts raw text characters into a hierarchy of typed objects. In JavaScript and TypeScript ecosystems, @babel/parser provides a battle-tested engine capable of parsing modern ECMAScript 2024 and TypeScript 5.7 syntax.

# Install parser and AST traversal libraries
npm install @babel/parser @babel/traverse @babel/types
npm install -D @types/babel__traverse @types/babel__core typescript tsx

Here is how to load and parse an in-memory TypeScript module:

// src/parser.ts
import { parse } from '@babel/parser';
import type { File } from '@babel/types';

export function parseSourceToAST(sourceCode: string): File {
  return parse(sourceCode, {
    sourceType: 'module',
    plugins: [
      'typescript',
      'jsx',
      'asyncGenerators',
      'dynamicImport',
      'objectRestSpread',
    ],
  });
}

Step 2: Walking the AST to Extract Scopes and Symbols

Once parsed into a tree, we walk the nodes using @babel/traverse. We do not want to pass 20,000 raw AST nodes to the model; that would waste context tokens. Instead, we extract high-signal structural metadata: declared functions, input arguments, return statements, external imports, and variable mutations.

// src/extractor.ts
import traverseModule from '@babel/traverse';
import type { File } from '@babel/types';

// Handle Babel CommonJS interop in ES modules
const traverse = (traverseModule as any).default || traverseModule;

export interface FunctionMetadata {
  name: string;
  startLine: number;
  endLine: number;
  params: string[];
  isAsync: boolean;
  calls: string[];
  mutations: string[];
  complexity: number;
}

export function extractCodeStructure(ast: File): {
  imports: Record<string, string[]>;
  functions: FunctionMetadata[];
} {
  const imports: Record<string, string[]> = {};
  const functions: FunctionMetadata[] = [];

  traverse(ast, {
    ImportDeclaration(path: any) {
      const source = path.node.source.value;
      const specifiers = path.node.specifiers.map((s: any) => s.local.name);
      imports[source] = specifiers;
    },

    FunctionDeclaration(path: any) {
      const funcName = path.node.id?.name || 'anonymous';
      const startLine = path.node.loc?.start.line || 0;
      const endLine = path.node.loc?.end.line || 0;
      const params = path.node.params.map((p: any) => p.name || p.type);
      const isAsync = path.node.async;

      const calls: string[] = [];
      const mutations: string[] = [];
      let complexity = 1; // Base cyclomatic complexity

      path.traverse({
        CallExpression(callPath: any) {
          const callee = callPath.node.callee;
          if (callee.name) {
            calls.push(callee.name);
          } else if (callee.property?.name) {
            calls.push(callee.property.name);
          }
        },
        AssignmentExpression(assignPath: any) {
          if (assignPath.node.left.name) {
            mutations.push(assignPath.node.left.name);
          }
        },
        IfStatement() { complexity++; },
        ForStatement() { complexity++; },
        WhileStatement() { complexity++; },
        ConditionalExpression() { complexity++; }, // Ternary ?:
      });

      functions.push({
        name: funcName,
        startLine,
        endLine,
        params,
        isAsync,
        calls: Array.from(new Set(calls)),
        mutations: Array.from(new Set(mutations)),
        complexity,
      });
    },
  });

  return { imports, functions };
}

Step 3: Enriching Code Embeddings with AST Metadata

Standard retrieval-augmented generation (RAG) splits text into arbitrary 500-token chunks. When a chunk boundary cuts through the middle of a conditional branch or for-loop, semantic integrity vanishes.

AST-aware chunking solves this by creating chunks on exact boundary nodes (individual functions or classes). We then prepend structured metadata before embedding:

// src/embedding-helper.ts
export function buildEnrichedChunk(
  filePath: string,
  func: FunctionMetadata,
  rawSnippet: string
): string {
  return `File: ${filePath}
Function: ${func.name}
Parameters: (${func.params.join(', ')})
Async: ${func.isAsync}
Cyclomatic Complexity: ${func.complexity}
External Calls: [${func.calls.join(', ')}]
Mutations: [${func.mutations.join(', ')}]

Implementation:
${rawSnippet}`;
}

When this enriched string is passed to an embedding model (like OpenAI text-embedding-3-small or local BGE-large-en), vector queries matching terms like "functions mutating state without database transaction" retrieve the exact function node with high cosine similarity.

Step 4: Feeding Enriched Structure to the LLM

Now we assemble the prompt. Instead of asking the model to deduce architecture from scratch, we hand it verified structural facts extracted directly from the parser.

// src/explainer.ts
import { parseSourceToAST } from './parser.js';
import { extractCodeStructure } from './extractor.js';

export function buildExplainerPrompt(
  filePath: string,
  sourceCode: string
): string {
  const ast = parseSourceToAST(sourceCode);
  const { imports, functions } = extractCodeStructure(ast);

  const metadataJson = JSON.stringify({ imports, functions }, null, 2);

  return `You are an expert compiler engineer explaining source code to a developer.
Use both the verified AST structural metadata and the raw source code below to provide an authoritative walkthrough.

### Verified AST Metadata:
${metadataJson}

### Source Code (${filePath}):
\`\`\`typescript
${sourceCode}
\`\`\`

Instructions:
1. State the purpose of each function, its cyclomatic complexity, and why it is significant.
2. Detail any mutable state changes identified in the AST.
3. Trace external dependencies listed in imports.
4. Call out edge cases such as unhandled promise rejections or unchecked null values.`;
}

Comparing Raw Prompting vs. AST-Guided Output

Consider a token refresh middleware that contains variable shadowing: an inner block declares const token that overrides an outer function parameter.

  • Raw Prompt Result: The model frequently confuses the inner token with the authorization header token, claiming that the header string is directly decoded without validation.
  • AST-Guided Result: The visitor detects two distinct lexical scopes. The model accurately notes: "Line 24 declares an inner lexical binding token that shadows the outer argument. The outer argument is preserved for telemetry logging."

Production Edge Cases to Handle

When running AST analysis in automated development pipelines, protect against three failure patterns:

  • Syntax Errors During Editing: Developers often run explainers while typing half-finished code. Wrap the parser in try/catch and fallback to tolerant parsing plugins (such as Babel's errorRecovery: true).
  • Dynamic Property Access: ASTs cannot determine values computed dynamically at runtime (e.g. target[computedKey]()). Mark these expressions explicitly as dynamic calls in your metadata output.
  • Macro and Decorator Overhead: Frameworks like NestJS, Angular, and MobX rely heavily on decorators. Ensure your parser configuration enables the decorators-legacy or modern decorators plugin to prevent parser crashes on annotations.

By shifting structural parsing work to a deterministic compiler and reserving the LLM for high-level semantic synthesis, you eliminate hallucinated explanations and deliver reliable code intelligence to your team.

Found this useful?
View all articles
Free Technical Interview Prep

Practicing for Engineering Interviews?

Skip the expensive coaching bootcamps and dry LeetCode memorization. Practice real production scenarios with instant turn-by-turn AI feedback on Frontend, Backend, System Design, and DSA.

Free Utilities

Recommended Developer Tools for this Topic

Explore all 25+ tools→

Keep Reading

Related Articles

Learn with Dropout Developer

Build real software with AI

Step-by-step learning paths, vibe coding tutorials, and certified developer programs designed for the modern engineer.