The Problem with Raw Code Prompting
Most developers assume AI code explanation tools work by feeding raw source strings directly into a large language model. You copy an unfamiliar 150-line utility function, send it to a chat prompt, and receive an explanation. When the snippet is trivial, this approach appears to work. On production codebases with complex lexical scopes, asynchronous callbacks, and cross-module dependencies, naive prompting breaks down.
Language models process text as linear sequences of sub-word tokens. Code is fundamentally non-linear: it is a graph of scopes, references, type assertions, and execution branches. When fed raw source text without structural scaffolding, models easily mistake shadowed variables for outer declarations, miss implicit closures, and hallucinate runtime control flows. Production-grade code explainers solve this by pairing Abstract Syntax Tree (AST) traversal with dense contextual embeddings.
The Architecture: From Lexical Tokens to Structural Intelligence
Before any prompt reaches an LLM, the explanation pipeline converts source code into a structured intermediate representation. The AST extracts exact symbol relationships, while vector embeddings locate related definitions across the repository.
Source Code (.ts / .js)
│
▼
[Lexer & Parser (@babel/parser)] ───► Generates Abstract Syntax Tree
│
▼
[AST Visitor (@babel/traverse)] ───► Extracts: Functions, Scopes, Cyclomatic Complexity
│
▼
[Vector Search (pgvector/LanceDB)] ──► Retrieves related schemas and imported interfaces
│
▼
[LLM Structured Prompt] ───► Generates line-accurate, hallucination-free explanation
Step 1: Parsing Source Code into an AST
An AST converts raw text characters into a hierarchy of typed objects. In JavaScript and TypeScript ecosystems, @babel/parser provides a battle-tested engine capable of parsing modern ECMAScript 2024 and TypeScript 5.7 syntax.
# Install parser and AST traversal libraries
npm install @babel/parser @babel/traverse @babel/types
npm install -D @types/babel__traverse @types/babel__core typescript tsx
Here is how to load and parse an in-memory TypeScript module:
// src/parser.ts
import { parse } from '@babel/parser';
import type { File } from '@babel/types';
export function parseSourceToAST(sourceCode: string): File {
return parse(sourceCode, {
sourceType: 'module',
plugins: [
'typescript',
'jsx',
'asyncGenerators',
'dynamicImport',
'objectRestSpread',
],
});
}
Step 2: Walking the AST to Extract Scopes and Symbols
Once parsed into a tree, we walk the nodes using @babel/traverse. We do not want to pass 20,000 raw AST nodes to the model; that would waste context tokens. Instead, we extract high-signal structural metadata: declared functions, input arguments, return statements, external imports, and variable mutations.
// src/extractor.ts
import traverseModule from '@babel/traverse';
import type { File } from '@babel/types';
// Handle Babel CommonJS interop in ES modules
const traverse = (traverseModule as any).default || traverseModule;
export interface FunctionMetadata {
name: string;
startLine: number;
endLine: number;
params: string[];
isAsync: boolean;
calls: string[];
mutations: string[];
complexity: number;
}
export function extractCodeStructure(ast: File): {
imports: Record<string, string[]>;
functions: FunctionMetadata[];
} {
const imports: Record<string, string[]> = {};
const functions: FunctionMetadata[] = [];
traverse(ast, {
ImportDeclaration(path: any) {
const source = path.node.source.value;
const specifiers = path.node.specifiers.map((s: any) => s.local.name);
imports[source] = specifiers;
},
FunctionDeclaration(path: any) {
const funcName = path.node.id?.name || 'anonymous';
const startLine = path.node.loc?.start.line || 0;
const endLine = path.node.loc?.end.line || 0;
const params = path.node.params.map((p: any) => p.name || p.type);
const isAsync = path.node.async;
const calls: string[] = [];
const mutations: string[] = [];
let complexity = 1; // Base cyclomatic complexity
path.traverse({
CallExpression(callPath: any) {
const callee = callPath.node.callee;
if (callee.name) {
calls.push(callee.name);
} else if (callee.property?.name) {
calls.push(callee.property.name);
}
},
AssignmentExpression(assignPath: any) {
if (assignPath.node.left.name) {
mutations.push(assignPath.node.left.name);
}
},
IfStatement() { complexity++; },
ForStatement() { complexity++; },
WhileStatement() { complexity++; },
ConditionalExpression() { complexity++; }, // Ternary ?:
});
functions.push({
name: funcName,
startLine,
endLine,
params,
isAsync,
calls: Array.from(new Set(calls)),
mutations: Array.from(new Set(mutations)),
complexity,
});
},
});
return { imports, functions };
}
Step 3: Enriching Code Embeddings with AST Metadata
Standard retrieval-augmented generation (RAG) splits text into arbitrary 500-token chunks. When a chunk boundary cuts through the middle of a conditional branch or for-loop, semantic integrity vanishes.
AST-aware chunking solves this by creating chunks on exact boundary nodes (individual functions or classes). We then prepend structured metadata before embedding:
// src/embedding-helper.ts
export function buildEnrichedChunk(
filePath: string,
func: FunctionMetadata,
rawSnippet: string
): string {
return `File: ${filePath}
Function: ${func.name}
Parameters: (${func.params.join(', ')})
Async: ${func.isAsync}
Cyclomatic Complexity: ${func.complexity}
External Calls: [${func.calls.join(', ')}]
Mutations: [${func.mutations.join(', ')}]
Implementation:
${rawSnippet}`;
}
When this enriched string is passed to an embedding model (like OpenAI text-embedding-3-small or local BGE-large-en), vector queries matching terms like "functions mutating state without database transaction" retrieve the exact function node with high cosine similarity.
Step 4: Feeding Enriched Structure to the LLM
Now we assemble the prompt. Instead of asking the model to deduce architecture from scratch, we hand it verified structural facts extracted directly from the parser.
// src/explainer.ts
import { parseSourceToAST } from './parser.js';
import { extractCodeStructure } from './extractor.js';
export function buildExplainerPrompt(
filePath: string,
sourceCode: string
): string {
const ast = parseSourceToAST(sourceCode);
const { imports, functions } = extractCodeStructure(ast);
const metadataJson = JSON.stringify({ imports, functions }, null, 2);
return `You are an expert compiler engineer explaining source code to a developer.
Use both the verified AST structural metadata and the raw source code below to provide an authoritative walkthrough.
### Verified AST Metadata:
${metadataJson}
### Source Code (${filePath}):
\`\`\`typescript
${sourceCode}
\`\`\`
Instructions:
1. State the purpose of each function, its cyclomatic complexity, and why it is significant.
2. Detail any mutable state changes identified in the AST.
3. Trace external dependencies listed in imports.
4. Call out edge cases such as unhandled promise rejections or unchecked null values.`;
}
Comparing Raw Prompting vs. AST-Guided Output
Consider a token refresh middleware that contains variable shadowing: an inner block declares const token that overrides an outer function parameter.
- Raw Prompt Result: The model frequently confuses the inner token with the authorization header token, claiming that the header string is directly decoded without validation.
- AST-Guided Result: The visitor detects two distinct lexical scopes. The model accurately notes: "Line 24 declares an inner lexical binding token that shadows the outer argument. The outer argument is preserved for telemetry logging."
Production Edge Cases to Handle
When running AST analysis in automated development pipelines, protect against three failure patterns:
- Syntax Errors During Editing: Developers often run explainers while typing half-finished code. Wrap the parser in try/catch and fallback to tolerant parsing plugins (such as Babel's
errorRecovery: true). - Dynamic Property Access: ASTs cannot determine values computed dynamically at runtime (e.g.
target[computedKey]()). Mark these expressions explicitly as dynamic calls in your metadata output. - Macro and Decorator Overhead: Frameworks like NestJS, Angular, and MobX rely heavily on decorators. Ensure your parser configuration enables the
decorators-legacyor moderndecoratorsplugin to prevent parser crashes on annotations.
By shifting structural parsing work to a deterministic compiler and reserving the LLM for high-level semantic synthesis, you eliminate hallucinated explanations and deliver reliable code intelligence to your team.
