TypeScript 98 ๐ท TypeScript Compiler Internals โ Overview
Every time you run tsc, your TypeScript code passes through a pipeline of transformations. The scanner breaks the source text into tokens. The parser assembles the tokens into an Abstract Syntax Tree. The binder connects declarations that refer to the same entity. The checker resolves types and reports semantic errors. The emitter produces JavaScript output. Each stage is a separate file in the compiler source, and each has a specific role that has not changed significantly since the compiler was first written in 2012 .
This chapter is an overview of that pipeline. It is not a guide to writing a compiler. It is a guide to understanding what happens when you press save in your editor or run tsc in CI, and why certain errors appear at certain stages. The internals are documented in the TypeScript wiki’s “Architectural Overview” page and in the community-maintained TypeScript Deep Dive book .
Key point: The TypeScript compiler is not a single program. It is a set of layers that can be used independently. The core pipeline (scanner, parser, binder, checker, emitter) is the compiler. The tsc command is a batch CLI that wraps the pipeline. The tsserver is a language server that exposes the pipeline through a JSON protocol for editors. The same pipeline powers both the command line and the IDE .
Why compiler internals matter
You do not need to know the compiler internals to write TypeScript. But knowing them explains behaviors that otherwise seem arbitrary.
The error location problem. When you see an error underlined in your editor, the position comes from a specific node in the syntax tree. The compiler tracks positions at the token level, and each node knows where it starts and ends. Understanding this explains why errors sometimes appear on the wrong line or point at the wrong token .
The incremental compile problem. The --watch mode and the language service do not re-run the full pipeline on every change. They patch the existing syntax tree and re-check only the affected files. This is why the editor is faster than running tsc from scratch. The pipeline is designed to be long-lived and incrementally updated .
The declaration file problem. When you import a library, TypeScript resolves the import to a .ts file first, then a .d.ts file. The pre-processor walks the import graph and builds the program from the ordered list of files. Understanding this explains why a missing .d.ts file produces a different error than a missing .ts file .
The performance problem. The TypeScript 7.0 rewrite in Go achieved 8xโ12x speedups by using native code, shared memory concurrency, and a more efficient traversal of the syntax graph. The pipeline stages are the same, but the implementation is different. Understanding the pipeline explains what can be optimized .
a. The Five Stages of the Pipeline
The TypeScript compiler processes source code in five stages. Each stage has a single responsibility and passes its output to the next stage.
Scanner. The scanner reads the source text character by character and produces a stream of tokens. A token is the smallest meaningful unit of the language: an identifier, a keyword, a number literal, a string literal, an operator, or a punctuation mark. The scanner also tracks trivia โ whitespace, comments, and conflict markers โ which are not part of the syntax tree but are preserved for use by the emitter and the language service .
Parser. The parser consumes the token stream and produces an Abstract Syntax Tree. The AST is a tree of nodes, where each node represents a construct in the language: a function declaration, an if statement, a binary expression, a type annotation. The parser follows the productions of the TypeScript grammar. It does not check types โ it only checks whether the code is syntactically valid .
Binder. The binder walks the AST and creates Symbols. A Symbol represents a named declaration: a variable, a function, a class, an interface, a module. When the same entity is declared multiple times โ two interfaces with the same name merging, a function and a namespace with the same name โ the binder connects those declarations to a single Symbol. The binder also handles scope resolution, determining which Symbol a name refers to in a given context .
Checker. The checker is the largest and most complex part of the compiler. It uses the AST and the Symbols to resolve types, check semantic operations, and produce diagnostics. The checker validates that a function’s return type matches its declared return type, that a variable is not used before it is declared, that an argument passed to a function matches the parameter’s type. It is also the source of most of the errors you see in your editor .
Emitter. The emitter takes the AST and the checker’s results and produces output. The output can be JavaScript (.js), declaration files (.d.ts), or source maps (.js.map). The emitter does not re-check types โ it trusts the checker’s output. If the checker produced no errors, the emitter generates JavaScript that reflects the type-erased source .
The pipeline is not linear in practice. The checker and the language service can query the AST and Symbols in any order. The emitter can be invoked without the checker for transpile-only tools like esbuild and SWC. The stages are composable, and different tools use different subsets of the pipeline .
b. The Supporting Layers
Beyond the core pipeline, the compiler has several supporting layers that make it usable in different contexts.
The Program. A Program is a collection of SourceFiles and a set of compiler options that represent a compilation unit. The Program is the main entry point to the type system and code generation. It owns the source files, the options, and the checker. When you run tsc, the Program is created from the files on the command line and their imports .
The CompilerHost. The CompilerHost is the interface through which the Program interacts with the operating system. It provides methods for reading files, checking whether files exist, and writing output. The Node.js implementation reads from the filesystem. The language service implementation reads from the editor’s in-memory buffers. This abstraction is what allows the same compiler to run in different environments .
The Language Service. The Language Service is an additional layer around the core pipeline, designed for editor-like applications. It provides completions, signature help, code formatting, rename, and other IDE features. The Language Service is designed to efficiently handle files changing over time within a long-lived compilation context. It does not re-run the full pipeline on every keystroke โ it patches the AST and re-checks only what is affected .
The Standalone Server (tsserver). The tsserver wraps the Language Service and exposes it through a JSON protocol. The editor communicates with the tsserver using this protocol. This is why VS Code can show TypeScript errors, provide completions, and perform renames without running tsc in the background .
The Pre-processor. The pre-processor figures out what files should be included in the compilation. It starts with the files passed on the command line, then follows import statements and /// <reference path=... /> tags to build the complete set of source files. When resolving imports, preference is given to .ts files over .d.ts files, ensuring the most up-to-date files are processed .
c. The Data Structures
The compiler is built on a small set of data structures that are used throughout the pipeline.
Node. A Node is the basic building block of the Abstract Syntax Tree. Each node represents a construct in the language: an identifier, a literal, a declaration, an expression, a statement. The node type is identified by the SyntaxKind enum. Nodes have a pos (the position in the file where the node starts) and an end (where it ends). The parser creates nodes; the binder and checker walk them .
SourceFile. A SourceFile is the AST of a single source file. It is itself a Node, and it provides additional interfaces for accessing the raw text of the file, the references in the file, the list of identifiers, and the mapping from a position in the file to a line and character number. The SourceFile is the unit that the parser produces and the checker consumes .
Symbol. A Symbol is a named declaration. Symbols are created by the binder. They connect declaration nodes in the AST to other declarations that contribute to the same entity. A class and a namespace with the same name share a Symbol. A function declaration and its overload signatures share a Symbol. The Symbol is the basic building block of the semantic system โ the checker reasons about the program in terms of Symbols, not raw AST nodes .
Type. A Type is the other half of the semantic system. Types can be named (a class, an interface) or anonymous (an object type, a union). The checker creates Types as it resolves the program. Every Symbol that has a type is associated with a Type object. The Type objects are interconnected โ a union type references its member types, a function type references its parameter and return types โ and the checker traverses this graph to perform type operations .
Signature. A Signature represents a call, construct, or index signature. A function has a call signature. A class constructor has a construct signature. An object with an index signature has an index signature. Signatures are associated with Types and are used by the checker to validate calls and assignments .
Program. The Program is the top-level structure. It owns the SourceFiles, the compiler options, and the checker. It is the entry point for type checking and code generation. When you create a Program, you provide a root set of files and a CompilerHost. The Program resolves the imports, creates SourceFiles, and makes the checker available .
The TypeScript 7.0 rewrite preserved these data structures but reimplemented them in Go. The scanner, parser, binder, checker, and emitter still exist, but they are now compiled to native code and can use shared memory concurrency. The performance gains come from the implementation, not from a change in the architecture .
Complete Example Session
This session demonstrates the pipeline by using the TypeScript Compiler API to create a Program, inspect the AST, and retrieve Symbols and Types.
// ============================================
// PART 1: THE SAMPLE CODE
// ============================================
// sample.ts
interface User {
id: string;
name: string;
}
function greet(user: User): string {
return `Hello, ${user.name}`;
}
const alice: User = { id: '1', name: 'Alice' };
greet(alice);
// ============================================
// PART 2: CREATE A PROGRAM
// ============================================
// inspect.ts
import * as ts from 'typescript';
const program = ts.createProgram(['sample.ts'], {
target: ts.ScriptTarget.ES2020,
module: ts.ModuleKind.ESNext,
strict: true,
});
// ============================================
// PART 3: GET THE SOURCE FILE (PARSER OUTPUT)
// ============================================
const sourceFile = program.getSourceFile('sample.ts')!;
console.log(sourceFile.fileName); // sample.ts
console.log(sourceFile.statements.length); // 3 statements
// ============================================
// PART 4: WALK THE AST
// ============================================
function walk(node: ts.Node, depth: number = 0) {
console.log(' '.repeat(depth) + ts.SyntaxKind[node.kind]);
node.forEachChild(child => walk(child, depth + 1));
}
walk(sourceFile);
// Output:
// SourceFile
// InterfaceDeclaration
// Identifier
// PropertySignature
// Identifier
// StringKeyword
// PropertySignature
// Identifier
// StringKeyword
// FunctionDeclaration
// Identifier
// Parameter
// Identifier
// TypeReference
// Identifier
// TypeReference
// Identifier
// Block
// ReturnStatement
// TemplateExpression
// ...
// VariableStatement
// ...
// ============================================
// PART 5: GET THE CHECKER (BINDER + CHECKER OUTPUT)
// ============================================
const checker = program.getTypeChecker();
// ============================================
// PART 6: FIND A SYMBOL
// ============================================
// Find the User interface
let userInterface: ts.InterfaceDeclaration | undefined;
sourceFile.forEachChild(node => {
if (ts.isInterfaceDeclaration(node) && node.name.text === 'User') {
userInterface = node;
}
});
if (userInterface) {
const symbol = checker.getSymbolAtLocation(userInterface.name);
console.log(symbol?.name); // User
}
// ============================================
// PART 7: GET THE TYPE OF A SYMBOL
// ============================================
if (userInterface) {
const symbol = checker.getSymbolAtLocation(userInterface.name)!;
const type = checker.getDeclaredTypeOfSymbol(symbol);
console.log(checker.typeToString(type));
// { id: string; name: string; }
}
// ============================================
// PART 8: CHECK FOR ERRORS
// ============================================
const diagnostics = ts.getPreEmitDiagnostics(program);
if (diagnostics.length === 0) {
console.log('No errors');
} else {
diagnostics.forEach(diag => {
const message = ts.flattenDiagnosticMessageText(diag.messageText, '\n');
console.log(`Error: ${message}`);
});
}
// ============================================
// PART 9: EMIT OUTPUT
// ============================================
const emitResult = program.emit();
console.log('Emitted files:', emitResult.emittedFiles);
// ============================================
// PART 10: THE FULL PIPELINE IN ONE CALL
// ============================================
// The tsc CLI does all of the above in one command:
// npx tsc --noEmit --strict sample.ts
The ten parts cover the sample code, creating a Program, getting the SourceFile (parser output), walking the AST, getting the checker (binder and checker output), finding a Symbol, getting the Type of a Symbol, checking for errors, emitting output, and the full pipeline in one call.
Quick Reference
The Five Core Stages
| Stage | Input | Output | File |
|---|---|---|---|
| Scanner | Source text | Token stream | scanner.ts |
| Parser | Token stream | AST | parser.ts |
| Binder | AST | Symbols | binder.ts |
| Checker | AST + Symbols | Type validation | checker.ts |
| Emitter | AST + Checker | JS / .d.ts / maps | emitter.ts |
The Supporting Layers
| Layer | Purpose |
|---|---|
| Program | Collection of SourceFiles and options |
| CompilerHost | OS abstraction (filesystem, editor) |
| Language Service | Editor features (completions, rename) |
| tsserver | JSON protocol for editor communication |
| Pre-processor | Resolves imports and builds the file list |
The Data Structures
| Structure | Purpose |
|---|---|
| Node | AST building block |
| SourceFile | AST of one file |
| Symbol | Named declaration |
| Type | Semantic type |
| Signature | Call/construct/index signature |
| Program | Top-level compilation unit |
The Compiler API Entry Points
| Function | Purpose |
|---|---|
ts.createProgram() | Create a Program |
program.getSourceFile() | Get the AST of a file |
program.getTypeChecker() | Get the checker |
checker.getSymbolAtLocation() | Find a Symbol |
checker.getTypeOfSymbolAtLocation() | Get a Type |
program.emit() | Generate output |
ts.getPreEmitDiagnostics() | Get all errors |
The SyntaxKind Enum
| Kind | Represents |
|---|---|
SourceFile | The root node of a file |
InterfaceDeclaration | An interface |
FunctionDeclaration | A function |
VariableStatement | A variable declaration |
Identifier | A name |
StringKeyword | The string type |
TypeReference | A reference to a named type |
Best Practices
โ Do This:
// Use the Compiler API to inspect the AST
const program = ts.createProgram(['file.ts'], {}); // โ
// Walk the AST with forEachChild
node.forEachChild(child => walk(child)); // โ
// Use the checker to resolve types
const type = checker.getTypeAtLocation(node); // โ
// Check for errors with getPreEmitDiagnostics
ts.getPreEmitDiagnostics(program); // โ
โ Don’t Do This:
// Don't use the compiler internals for transpilation
// Use esbuild or SWC for speed. // โ
// Don't rely on the AST shape staying stable
// The API is not versioned like the language. // โ
// Don't build a language service without understanding
// the incremental compilation model. // โ
Common Pitfalls
| Pitfall | Why It Happens | Fix |
|---|---|---|
Cannot find module 'typescript' | Compiler API not imported | import * as ts from 'typescript' |
| AST node has no type | Checker not initialized | Use program.getTypeChecker() |
| Diagnostics empty | noEmit not set | Use getPreEmitDiagnostics |
| Performance issues | Re-creating Program | Reuse the Program, patch files |
| API version mismatch | Using internal APIs | Use public APIs only |
Real-World Examples
1. Create a Program
const program = ts.createProgram(['file.ts'], { strict: true });
2. Get a SourceFile
const sourceFile = program.getSourceFile('file.ts')!;
3. Walk the AST
sourceFile.forEachChild(node => console.log(ts.SyntaxKind[node.kind]));
4. Get the Checker
const checker = program.getTypeChecker();
5. Find a Symbol
const symbol = checker.getSymbolAtLocation(node.name);
6. Get a Type
const type = checker.getTypeAtLocation(node);
7. Get Diagnostics
const diagnostics = ts.getPreEmitDiagnostics(program);
8. Emit Output
const emitResult = program.emit();
9. Incremental Compilation
const program = ts.createProgram(['file.ts'], { incremental: true });
10. Language Service
const service = ts.createLanguageService(host);
Visual
The Compiler Pipeline
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ TYPESCRIPT COMPILER PIPELINE โ
โ โ
โ Source text โ
โ โ โ
โ โผ โ
โ Scanner โโ> Token stream โ
โ โ โ
โ โผ โ
โ Parser โโ> AST โ
โ โ โ
โ โผ โ
โ Binder โโ> Symbols โ
โ โ โ
โ โผ โ
โ Checker โโ> Type validation + diagnostics โ
โ โ โ
โ โผ โ
โ Emitter โโ> JS / .d.ts / source maps โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Data Flow
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ DATA FLOW โ
โ โ
โ SourceCode ~~ scanner ~~> Token Stream โ
โ โ โ
โ โผ โ
โ Token Stream ~~ parser ~~> AST โ
โ โ โ
โ โผ โ
โ AST ~~ binder ~~> Symbols โ
โ โ โ
โ โผ โ
โ AST + Symbols ~~ checker ~~> Type Validationโ
โ โ โ
โ โผ โ
โ AST + Checker ~~ emitter ~~> JS โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The Supporting Layers
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SUPPORTING LAYERS โ
โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ tsserver (JSON protocol) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ
โ โ Language Service (editor features) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ
โ โ Core Compiler Pipeline โ โ
โ โ (scanner, parser, binder, checker, โ โ
โ โ emitter) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ
โ โ CompilerHost (OS abstraction) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค โ
โ โ Program (SourceFiles + options) โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
The TypeScript 7.0 Native Port
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ TYPESCRIPT 6 vs 7 โ
โ โ
โ TypeScript 6: โ
โ Written in TypeScript โ
โ Runs on V8 (Node.js) โ
โ Single-threaded โ
โ โ
โ TypeScript 7: โ
โ Written in Go โ
โ Compiled to native code โ
โ Shared memory multithreading โ
โ --checkers and --builders flags โ
โ --singleThreaded for constrained envs โ
โ โ
โ Pipeline stages are the same. โ
โ Implementation is different. โ
โ Speedup: 8xโ12x on full builds. โ
โ โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Summary
| Item | Value |
|---|---|
| Scanner | Source text โ tokens |
| Parser | Tokens โ AST |
| Binder | AST โ Symbols |
| Checker | AST + Symbols โ type validation |
| Emitter | AST + Checker โ JS / .d.ts |
| Program | Collection of SourceFiles + options |
| CompilerHost | OS abstraction |
| Language Service | Editor features |
| tsserver | JSON protocol for editors |
| Node | AST building block |
| Symbol | Named declaration |
| Type | Semantic type |
| TypeScript 7.0 | Go rewrite, 8xโ12x faster |
Key takeaways:
- The compiler is a pipeline of five stages: scanner, parser, binder, checker, emitter. Each stage has a single responsibility and passes its output to the next. The scanner produces tokens, the parser produces the AST, the binder produces Symbols, the checker validates types, and the emitter produces output .
- The Program is the top-level compilation unit. It owns the SourceFiles, the compiler options, and the checker. It is the entry point for both type checking and code generation. The
ts.createProgram()API creates a Program from a list of files . - The CompilerHost abstracts the operating system. It provides methods for reading files and writing output. The Node.js implementation reads from the filesystem; the language service implementation reads from the editor’s in-memory buffers. This abstraction allows the same compiler to run in different environments .
- The Language Service and tsserver are layers on top of the core pipeline. They expose the compiler’s features โ completions, signature help, rename, diagnostics โ to editors through a JSON protocol. This is how VS Code shows TypeScript errors without running
tscin the background . - The AST is built from Nodes; the semantic system is built from Symbols and Types. The binder creates Symbols from the AST, connecting declarations that refer to the same entity. The checker creates Types from the Symbols and validates the program’s semantics .
- TypeScript 7.0 rewrote the compiler in Go. The pipeline stages are the same, but the implementation is native code with shared memory multithreading. The rewrite achieved 8xโ12x speedups on full builds while preserving full type checking .
- The compiler internals explain the editor’s behavior. The incremental compilation model, the position tracking, and the import resolution are all implemented in the pipeline. Understanding them makes the editor’s behavior predictable rather than mysterious .
Remember: The TypeScript compiler is not a black box. It is a pipeline of five stages that transform your source code into JavaScript. The scanner tokenizes, the parser builds the tree, the binder connects declarations, the checker validates types, and the emitter produces output. The Program, CompilerHost, Language Service, and tsserver are layers that make the pipeline usable in different contexts. The TypeScript 7.0 rewrite changed the implementation language to Go for speed, but the architecture is the same. Understanding the pipeline explains why errors appear where they do, why the editor is faster than tsc, and why the compiler can be extended by tools that use the same APIs.
Stop using slow, ad-bloated tool sites! ๐คฎ
๐ Search “KandZ Tools” on Google to use many professional utilities for free.
KandZ.me is the ultimate minimalist hub for:
โ
Finance (Mortgage, Interest, Inflation)
โ
Tech (Base64, JSON, Dev Suite, IP)
โ
Health (BMI, BMR, TDEE)
โ
Productivity (Timer, Workspace, QR)
โก๏ธ Fast & Private
๐ No data leaves your device
๐ 100% Free
๐ Use it now: https://tools.kandz.me
๐ Bookmark itโyouโll need it later!