Skip to main content
Docs menu

HTML filters

Last updated Edit this page View as Markdown

HTML optimization filters in mod_pagespeed 2.1: collapse whitespace, strip comments, elide attributes, DNS prefetch, and resource preload hints.

On this page

Overview

mod_pagespeed 2.1 includes filters that optimize HTML structure, cut unnecessary bytes, and add performance hints. Two of them (add_head, convert_meta_tags) are CoreFilters and run by default; the rest are opt-in — see Filter selection to turn them on. For resources rather than markup, see the CSS filters, JavaScript filters, and Image filters pages.

Quick reference

FilterCoreDescriptionSafe
add_headYesAdds <head> if missingYes
convert_meta_tagsYesConverts meta http-equiv to headersYes
collapse_whitespaceNoRemoves excess whitespaceGenerally safe
remove_commentsNoStrips HTML commentsGenerally safe
elide_attributesNoRemoves default-value attributesGenerally safe
remove_quotesNoRemoves unnecessary attribute quotesGenerally safe
trim_urlsNoShortens absolute URLs to relativeGenerally safe
combine_headsNoMerges multiple <head> elementsGenerally safe
pedanticNoAdds type attributes for HTML4Generally safe
insert_dns_prefetchNoAdds DNS prefetch hintsGenerally safe
hint_preload_subresourcesNoAdds preload headersGenerally safe
add_instrumentationNoInjects page load timing JSTest first
add_base_tagNoAdds a <base href> for the pageTest first
add_idsNoAdds ids to elements without oneTest first
insert_amp_linkNoAdds a <link rel=amphtml>Generally safe
debugNoExplains rewriting in commentsNot for production
decode_rewritten_urlsNoRestores original resource URLsTest first
compute_statisticsNoHTML statistics for the consoleTest first
experiment_http2NoHTTP/2 features in developmentExperimental
insert_gaNoInserts the retired ga.js snippetDeprecated

IIS syntax

On IIS, use the same filter names with the pagespeed prefix in pagespeed.config (no semicolons):

pagespeed EnableFilters collapse_whitespace,remove_comments

See IIS configuration for the full file format reference.

CoreFilters

add_head

Full guide →

Adds a <head> element if the HTML lacks one. Several other filters inject content into <head>, so this filter ensures one exists. Runs automatically as a CoreFilter.

convert_meta_tags

Full guide →

Reads <meta http-equiv="Content-Type"> and similar tags and adds corresponding HTTP response headers. This helps browsers discover the content type and character encoding earlier in the response. Runs automatically as a CoreFilter.

HTML minification filters

These filters reduce HTML payload size by removing unnecessary bytes.

collapse_whitespace

Full guide →

What it does

collapse_whitespace shrinks the HTML payload by folding every run of whitespace (spaces, tabs, carriage returns, newlines) down to a single character. The markup keeps its structure; only the formatting bytes go. Live demo: collapse_whitespace.

<!-- before -->
<ul class="nav">
    <li>  <a href="/a">Alpha</a>  </li>
    <li>  <a href="/b">Beta</a>   </li>
</ul>

<!-- after: each run that held a newline keeps one newline -->
<ul class="nav">
<li> <a href="/a">Alpha</a> </li>
<li> <a href="/b">Beta</a> </li>
</ul>

When it helps and when it does not

The saving scales with how much pretty-printing the templates do: indented server-side templates carry many collapsible bytes, already-compacted HTML almost none. Gzip and Brotli compress whitespace runs well, so the on-the-wire saving is much smaller than the source saving suggests. The filter earns most where HTML is served uncompressed, and least on an already minified template. Not a CoreFilter; enable it by name.

How it decides

A whitespace run that contains a newline collapses to a newline, any other run to a single space; a run never collapses to nothing, so inline elements keep their separating space and text layout is preserved. Content inside pre, code, script, style, and textarea is never touched, because whitespace is meaningful there.

Risks

  • Pages whose rendering depends on whitespace-sensitive CSS, such as white-space: pre on ordinary elements, can change appearance; the filter sees the markup, not the stylesheet. Test such pages before enabling.
  • Verify with the X-Mod-Pagespeed response header and a ?PageSpeedFilters=-collapse_whitespace comparison; Is it working? has the steps.

Configuration

# Apache
ModPagespeedEnableFilters collapse_whitespace
# Nginx
pagespeed EnableFilters collapse_whitespace;

remove_comments

Full guide →

What it does

remove_comments deletes HTML comments (<!-- ... -->) from the served page, cutting bytes that only developers read. The document tree is unchanged apart from the removed nodes. Live demo: remove_comments.

<!-- before -->
<!-- TODO: replace with the component version -->
<div class="banner">Sale ends Friday</div>

<!-- after -->
<div class="banner">Sale ends Friday</div>

When it helps and when it does not

It helps on pages whose templates carry heavy commentary: build markers, TODO notes, and section labels add up on large documents. It does nothing for pages that ship few comments, and it must not remove comments that carry a function. Comments you are obliged or willing to ship stay behind a RetainComment wildcard, and IE conditional comments (<!--[if IE]> ... <![endif]-->) are parsed as directives rather than comments, so they always survive.

How it decides

Every comment node is dropped unless its text matches one of the configured RetainComment wildcard patterns. There is no size threshold and no content analysis beyond the pattern match.

Risks

  • Copyright or license notices that must ship with the page need a RetainComment entry before the filter goes on.
  • A rare third-party snippet that reads the page’s own comments breaks; retain its marker comment or disable the filter for that path. Verify with the X-Mod-Pagespeed header and ?PageSpeedFilters=-remove_comments; Is it working? has the steps.

Configuration

# Apache
ModPagespeedEnableFilters remove_comments
ModPagespeedRetainComment "*copyright*"
# Nginx
pagespeed EnableFilters remove_comments;
pagespeed RetainComment "*copyright*";

elide_attributes

Full guide →

Removes HTML attributes that are set to their default values. For example, <form method="get"> becomes <form> because get is the default method.

remove_quotes

Full guide →

Removes unnecessary quotation marks around HTML attribute values when the value contains no special characters. Saves a few bytes per attribute.

trim_urls

Full guide →

What it does

trim_urls shortens URLs inside the page by stripping the parts that repeat the page’s own origin. An absolute URL whose scheme, host and port match the page loses its origin, and the page’s own directory prefix is trimmed too; URLs on any other origin, including the same host over another scheme, are left untouched. left_trim_urls is an accepted alternate spelling of the same filter; both names switch on the same code. Live demo: trim_urls.

<!-- page: https://example.com/shop/ -->
<!-- before -->
<a href="https://example.com/shop/cart">Cart</a>
<img src="https://example.com/img/logo.png" />

<!-- after -->
<a href="cart">Cart</a>
<img src="/img/logo.png" />

When it helps and when it does not

It saves a few bytes per URL on pages dense with same-origin absolute links, which is typical of CMS output that expands every URL in full; under gzip or Brotli the saving shrinks further, since the repeated origin strings compress well. It does not help pages that already use relative URLs throughout. Keep it off for HTML that lives beyond its origin: a saved page, an emailed copy, or markup served under a second domain resolves relative URLs against the wrong base.

How it decides

Each URL-valued attribute is resolved against the page’s base URL, and a <base> tag wins when present and is never itself rewritten. Only what matches gets trimmed: a full origin match drops the origin, a path under the page’s own directory then drops that directory as well, and a URL on another origin keeps its full form. A trim is kept only when the shorter URL resolves back to exactly the original one.

Risks

  • Serving the same cached HTML from multiple domains, or any flow that detaches the markup from its origin, turns the trimming into broken links; disable trim_urls for that content.
  • Verify with the X-Mod-Pagespeed response header and a ?PageSpeedFilters=-trim_urls comparison; Is it working? has the steps.

Configuration

# Apache
ModPagespeedEnableFilters trim_urls
# Nginx
pagespeed EnableFilters trim_urls;

Structural filters

add_base_tag

Full guide →

Adds a <base href="…"> element with the page’s own URL to <head>, so relative URLs in the page resolve against the URL the module rewrote them for. Use it when the HTML is served at a URL other than the one it was authored for (proxy setups, MapProxyDomain). Risk: a page that already relies on a different base, or on the absence of one, resolves its relative links differently; test navigation and form actions.

pagespeed EnableFilters add_base_tag;

add_ids

Full guide →

Adds an id attribute to elements that have none, so that beacon-driven filters can refer to individual elements across page loads. Rarely needed on its own: the filters that need ids enable it themselves. Risk: scripts or styles that count on the exact set of ids in the page see extra ones.

pagespeed EnableFilters add_ids;

combine_heads

Full guide →

What it does

combine_heads folds extra <head> elements into the first one. When the module’s HTML parse closes a second or later <head>, the contents of that element move into the first head and the emptied element disappears, so the served document carries exactly one head with everything in it. Not a core filter; enable it by name. Live demo: combine_heads.

<!-- before: two fragments, each contributing its own head -->
<head>
  <title>Report</title>
</head>
...the first fragment's content...
<head>
  <link rel="stylesheet" href="/css/part2.css" />
</head>
...the second fragment's content...

<!-- after: one head, both fragments' head content in it -->
<head>
  <title>Report</title>
  <link rel="stylesheet" href="/css/part2.css" />
</head>
...the first fragment's content... ...the second fragment's content...

When it helps and when it does not

It helps on pages assembled from fragments that each ship a complete document skeleton: server-side includes that embed whole HTML files, portal or search pages that splice in results blocks with their own <head>, scraped or migrated content pasted with its original head. A browser’s own parser discards the stray <head> tags and keeps the content, so such pages mostly render by accident; the merged output is predictable and gives the module’s head-injecting filters one place to work. On a normal page with a single head the filter changes nothing.

How it decides

The first <head> in the document is the survivor. Each later </head> moves its children into that first head, in document order. The merge cannot cross a flush boundary: when the server streams the page in chunks, heads that close in a later chunk than the first head stay as they are, because content already flushed to the browser cannot be restructured. The filter shares its implementation with add_head: the same pass that inserts a missing <head> performs the merge, so enabling combine_heads also brings add_head’s behavior of inserting a head into documents that have none.

Risks

  • Content ordering inside the merged head follows the fragments, not any intent the fragments had; a fragment whose styles or scripts assumed their original position relative to other elements should be checked after enabling.
  • Verify with the X-Mod-Pagespeed response header and a ?PageSpeedFilters=-combine_heads comparison; Is it working? has the steps.

Configuration

# Apache
ModPagespeedEnableFilters combine_heads
# Nginx
pagespeed EnableFilters combine_heads;

pedantic

Full guide →

Adds type="text/javascript" and type="text/css" attributes to <script> and <style> elements. This satisfies HTML4 validators. Not needed for HTML5, where these types are the defaults.

Performance hint filters

insert_dns_prefetch

Full guide →

What it does

insert_dns_prefetch adds connection warm-up hints for the origins of resources referenced in the body that the head does not already reference. A <link rel="dns-prefetch"> hint starts the DNS lookup early, while the browser is still busy with the HTML; a <link rel="preconnect"> hint goes further and opens the connection, TCP and for HTTPS also TLS, before the resource tag is even seen. Live demo: insert_dns_prefetch.

<!-- inserted into <head> -->
<link rel="preconnect" href="https://fonts.examplecdn.com" />
<link rel="dns-prefetch" href="//analytics.example.com" />

When it helps and when it does not

It helps when a page pulls from a few stable third-party origins, such as font CDNs or analytics hosts: the lookup and handshake then overlap with the HTML download instead of starting when the resource is discovered. It does nothing for same-origin resources, since the connection to the page’s own origin is already open. It also does little when the set of third-party domains churns between page views, because the hints are learned from earlier rewrites of the page and may name domains it no longer uses.

How it decides

The filter records which origins a page’s body resources come from, in order of first appearance. Hints are emitted once the number of hinted domains changes by at most two between rewrites; the check counts domains, so a page whose domains change but whose count holds still gets hints for the stored list. A page gets at most eight hints in total; the first two domains get preconnect, the rest dns-prefetch. Early views contribute data and get nothing; later views get the hints. Domains the author already hinted or referenced in the head are not duplicated.

Risks

  • Every hint costs the browser work, and a preconnect costs an open connection held for an origin the visitor might not need; the built-in caps of eight hints, at most two of them preconnects, bound that overhead.
  • Verify with the X-Mod-Pagespeed response header and a ?PageSpeedFilters=-insert_dns_prefetch comparison; Is it working? has the steps.

Configuration

# Apache
ModPagespeedEnableFilters insert_dns_prefetch
# Nginx
pagespeed EnableFilters insert_dns_prefetch;

hint_preload_subresources

Full guide →

Adds Link: rel=preload HTTP headers for CSS and JavaScript files discovered on previous visits to the same page. Uses the beacon system to collect resource data, so it becomes effective after the first page view.

Since v1.15.0+r21, <script type="module"> subresources are hinted with rel=modulepreload in the Link response header instead of rel=preload; modules carrying integrity or crossorigin="use-credentials" are left unhinted.

insert_speculation_rules

Full guide →

Injects a same-origin prefetch <script type="speculationrules"> block so that supporting browsers prefetch a link as the visitor starts interacting with it; browsers without speculation-rules support ignore the tag. Only same-origin links are eligible.

The filter stands down in several cases rather than injecting a ruleset that would be wasted or unsafe: it backs off when the page already carries its own speculation ruleset, when a Content-Security-Policy forbids inline scripts, on non-200 responses, on cookie-setting responses, on no-store responses, and on AMP documents.

The injected ruleset is fixed: one rule that prefetches same-origin document URLs (href_matches: "/*") with eagerness: moderate, meaning the browser applies its own heuristic for when a prefetch is worth starting rather than prefetching every link on sight. There is no prerender and nothing cross-origin. The script is placed at the end of <body>. Requests that negotiate markdown (Accept: text/markdown), such as AI-agent fetches, are also left clean: they run no scripts and get no benefit, so the tag stays out of the variant those caches key on under Vary: Accept.

Speculative prefetch spends origin bandwidth on navigations that may never happen, so treat it as a trade-off and test it against your own traffic. This filter is opt-in and is not part of any rewrite level. Enable it by name:

Apache:

ModPagespeedEnableFilters insert_speculation_rules

nginx:

pagespeed EnableFilters insert_speculation_rules;

Full guide →

Adds a <link rel="amphtml"> to <head> pointing at the page’s AMP version, built from the AmpLinkPattern directive. Use it when you publish AMP pages at a predictable URL pattern and want every canonical page to announce its AMP twin. Risk: a pattern that produces URLs that do not exist advertises broken AMP pages to crawlers.

pagespeed AmpLinkPattern "https://amp.example.com${url}";
pagespeed EnableFilters insert_amp_link;

Debugging and measurement filters

debug

Full guide →

Annotates the page with HTML comments that say which filters ran and why a resource was or was not rewritten (for example why an image was not inlined or a stylesheet not combined). Enable it per request with ?PageSpeedFilters=+debug while troubleshooting instead of in the configuration: it exposes internals, enlarges every page, and is not meant for production traffic.

https://www.example.com/?PageSpeedFilters=+debug

decode_rewritten_urls

Full guide →

Turns .pagespeed. resource URLs in the page back into the original resource URLs, undoing the URL rewriting of the other filters. Useful in a proxy chain or when debugging what the page referenced before optimization. Risk: the page then references unoptimized resources, which defeats the filters that depend on rewritten URLs.

pagespeed EnableFilters decode_rewritten_urls;

compute_statistics

Full guide →

Computes statistics about the HTML (element counts and sizes) for the admin console. It adds a parsing pass per page and changes nothing in the output; enable it while you need the numbers.

pagespeed EnableFilters compute_statistics;

experiment_http2

Full guide →

Switches on HTTP/2-specific behavior that is still in development. Experimental: what it does can change between releases, and it is not covered by the compatibility promises of the other filters.

pagespeed EnableFilters experiment_http2;

Analytics filters

insert_ga

Full guide →

Inserts the Google Analytics snippet for the account in AnalyticsID into every page. Deprecated: the snippet it inserts is the retired ga.js, and AnalyticsID itself only targets Universal Analytics, which was discontinued. Add your analytics in your templates instead.

Deprecated and dangerous filters

The names below are still accepted so that an existing configuration keeps loading, with a warning. The deprecated ones do nothing at all; fix_reflows and mobilize are in the dangerous set, which RewriteLevel AllFilters never enables, and exist for experiments rather than production. Remove them from your configuration.

cache_partial_html

Full guide →

Deprecated no-op.

defer_iframe

Full guide →

Deprecated no-op: iframe deferral is built into defer_javascript; enabling this name alone never did anything.

div_structure

Full guide →

Deprecated no-op.

explicit_close_tags

Full guide →

Deprecated no-op.

flush_subresources

Full guide →

Deprecated no-op.

fix_reflows

Full guide →

Experimental fix for layout reflows caused by deferred JavaScript. In the dangerous set; not for production.

mobilize

Full guide →

The retired page-mobilization experiment. In the dangerous set; not for production.

mobilize_precompute

Full guide →

Deprecated no-op.

split_html

Full guide →

Deprecated no-op.

split_html_helper

Full guide →

Deprecated no-op.

add_instrumentation

Full guide →

Injects JavaScript that measures page load time and reports it back to the mod_pagespeed statistics system via the beacon endpoint (/mod_pagespeed_beacon or /ngx_pagespeed_beacon). Enable this filter to get client-side performance data in the admin console histograms.

Test this filter before deploying to production. The injected JavaScript adds a small overhead and sends beacon requests on every page load.

Since v1.15.0+r18, with HonorCsp (default on) nothing is injected on pages whose Content-Security-Policy disallows the instrumentation script — no beacon fires for those pages, so they contribute no data to the console histograms.

See also

Search