HTML Parsing: How a Byte Stream Becomes a DOM

The HTML parser is the component every page waits on. This module follows bytes through tokenization and tree construction, then shows what interrupts that walk: a <script> with no async or defer stops the tokenizer mid-document — and that stop is what creates the preload scanner that races ahead to fetch markup the parser has not reached — while a stylesheet decides whether anything may paint. The lesson is that parse order is execution order, and where you put a tag is a scheduling decision.

Launch Simulator →

What you'll explore

  1. 01

    Tokenization

    The tokenizer is a state machine that consumes characters and emits tokens — start tags, end tags, characters, comments, DOCTYPE. It is not a parser in the grammar sense: HTML has no invalid input, only input that produces surprising tokens. Blink runs the tokenizer over a SegmentedString, which is why the parser can be handed the document in network-sized chunks and resume mid-token when the next chunk arrives. Document size matters here in the most literal way: tokenization cost is linear in bytes, and it is the one cost no tag placement can remove.

  2. 02

    Tree construction

    Tokens are fed one at a time to the tree builder, which maintains an insertion mode and a stack of open elements. This is where the specification's error recovery lives — an unclosed <p>, a <table> containing stray text, a </div> with nothing to close are all resolved here rather than rejected. Tree construction is where the DOM you inspect in devtools comes from, and why it frequently does not match the markup you wrote.

  3. 03

    Render blocking

    Two different things get called blocking and they block different things. A classic <script src> blocks the PARSER: no further tokens are processed until it has been fetched and executed, because the script may call document.write and change the bytes still to be parsed. A stylesheet blocks PAINT: the parser keeps building the DOM, but nothing is allowed on screen until the sheet has loaded, so that the first paint is not a flash of unstyled content. That second effect has a sharp boundary in the source — a stylesheet counts as render-blocking only while no <body> element exists yet, so a sheet linked from the middle of the body does not hold back paint at all. Both effects are visible in this module, and confusing them is the most common cause of a misdiagnosed slow page.

  4. 04

    Preload scanning

    While the main parser is stopped on a script, a second lightweight scanner keeps reading ahead through the buffered markup looking only for things worth fetching — src, href, imagesrcset. It builds no nodes and runs no script; it exists purely to get requests onto the network earlier than the parser could ask for them. This is why a blocking script in <head> costs less than the arithmetic suggests, and why the one thing that truly defeats the scanner is markup that does not exist yet: URLs assembled by script at runtime.

  5. 05

    DOMContentLoaded vs load

    DOMContentLoaded fires when the parser has finished the document and every deferred script has run — it says the DOM is complete, nothing about images or stylesheets. The load event waits for subresources as well. The gap between them is where most of a page's wall-clock time hides, and which of your scripts land before or after DOMContentLoaded is decided entirely by whether you wrote async, defer, or neither.

Validated against Chromium / Blink HTML parser (third_party/blink/renderer/core) chromium-154.0.8037.57 — third_party/blink/renderer/core/html/parser/html_tokenizer.cc (tokenizer state machine) and friends. Every simulated behavior cites a real source function, and a server-side correlation engine computes how each layer's state cascades into the next. See the methodology →

This module is part of zerohop Pro — it unlocks alongside the full kata library, progress tracking, and weak-area analysis.