GCVE Workshop - 22 September 2026 (14:00-18:00), Luxembourg Before The Vulnopticon Conference - Registration

GHSA-29PJ-957V-52MC

Vulnerability from github – Published: 2026-08-06 20:30 – Updated: 2026-08-06 20:30
VLAI
Summary
league/commonmark: AttributesExtension href/src unsafe-link filter bypass via embedded control bytes
Details

## Summary

The AttributesExtension's href/src unsafe-link filter (AttributesHelper::filterAttributes()) can be bypassed by embedding control bytes in a javascript: URL that browsers discard before parsing the scheme. Two variants:

  • Tab/newline inside the scheme — a literal ASCII TAB (0x09), CR (0x0D), or LF (0x0A), e.g. java<TAB>script:alert(1). Per the WHATWG URL Standard's "basic URL parser" step 3, browsers "remove all ASCII tab or newline from input".
  • Leading C0 controls — e.g. <0x01>javascript:alert(1). Per step 1 of the same algorithm, browsers remove any leading or trailing C0 control or space. (A leading space alone does not bypass, because parseAttributes() already trim()s the value; other C0 bytes are not trimmed.)

The filter is a literal anchored-prefix regex (RegexHelper::isLinkPotentiallyUnsafe() / REGEX_UNSAFE_PROTOCOL) that matches neither obfuscated form, so in both cases the browser still executes javascript:alert(1).

This is confirmed reproducible even with allow_unsafe_links => false set — i.e. even applications that have followed the library's own documented hardening guidance for untrusted input remain exploitable.

This is a sibling gap in the same defense that CVE-2025-46734 (GHSA-3527-qv2q-pfvx) fixed in v2.7.0 — that fix made href/src respect allow_unsafe_links, but did not normalize control bytes before checking, so these obfuscation techniques were never covered.

Vulnerability

Files: - src/Util/RegexHelper.php:69 (REGEX_UNSAFE_PROTOCOL), :239-242 (isLinkPotentiallyUnsafe()) - src/Extension/Attributes/Util/AttributesHelper.php:149-179 (filterAttributes())

CWE: CWE-79 (Improper Neutralization of Input During Web Page Generation / XSS) — primary - CWE-692 (Incomplete Denylist to Cross-Site Scripting) — the anchored-prefix denylist in REGEX_UNSAFE_PROTOCOL is incomplete. This is a composite of CWE-184 and CWE-79, so it captures the full "incomplete denylist → XSS" chain on its own. - CWE-86 (Improper Neutralization of Invalid Characters in Identifiers in Web Pages) — the specific evasion technique: control bytes embedded within the URI scheme identifier, which the browser strips before resolving it.

Root Cause

// src/Util/RegexHelper.php
public const REGEX_UNSAFE_PROTOCOL = '/^(?:javascript|vbscript|file|data):/i';

public static function isLinkPotentiallyUnsafe(string $url): bool
{
    return \preg_match(self::REGEX_UNSAFE_PROTOCOL, $url) !== 0 && \preg_match(self::REGEX_SAFE_DATA_PROTOCOL, $url) === 0;
}

// src/Extension/Attributes/Util/AttributesHelper.php
foreach ($attributes as $name => $value) {
    $attrNameLower = \strtolower($name);
    if (! $allowUnsafeLinks && ($attrNameLower === 'href' || $attrNameLower === 'src') && \is_string($value) && RegexHelper::isLinkPotentiallyUnsafe($value)) {
        unset($attributes[$name]);
        continue;
    }
    ...

The Attributes extension's own quote-value grammar (PARTIAL_DOUBLEQUOTEDVALUE = '"[^"]*"') accepts any byte except " inside quotes, including raw tab/CR/LF and other C0 controls, and parseAttributes() only trim()s (leading/trailing, and only the default charlist " \t\n\r\0\x0B" — so a leading \x01 survives). Critically, the core Markdown link-destination path (LinkParserHelperUrlEncoder::unescapeAndEncode()) percent-encodes every control byte before this same safety check ever runs — but the Attributes extension's href/src handling has no equivalent normalization step, so the raw control byte reaches both the check and the final HTML output (Xml::escape() only escapes & < > " ', not tab/CR/LF, since they're legal bytes inside an HTML attribute).

Attack Scenario

  1. An application enables the (commonly-used) AttributesExtension and sets allow_unsafe_links => false — the project's own documented hardening step for untrusted input.
  2. An attacker submits Markdown: [Click me](javascript:alert(0)){href="java<TAB>script:alert(document.cookie)"} (TAB is one literal 0x09 byte).
  3. The library emits <a href="java<TAB>script:alert(document.cookie)">Click me</a>isLinkPotentiallyUnsafe() doesn't match the tab-split scheme, so the filter takes no action.
  4. A victim viewing/clicking the link has the browser strip the embedded TAB and execute javascript:alert(document.cookie) in the victim's session — stored XSS, cookie theft, account takeover potential.

Why the payload needs an unsafe core destination. Step 2 above deliberately uses [Click me](javascript:alert(0)) rather than a normal link. LinkRenderer overwrites attrs['href'] with the node's own URL unless that URL is itself judged unsafe — so [x](https://example.com){href="java<TAB>script:..."} renders the harmless href="https://example.com", and an empty destination [x](){href="..."} renders href="". The attacker therefore supplies a core destination that the filter does catch, which suppresses the overwrite and lets the attribute-supplied href reach the final tag. This is no obstacle in practice — the attacker writes the entire Markdown document.

Two related forms that are not exploitable, noted so the fix isn't over-scoped:

  • Attaching the attribute to a non-link block — hi {href="java<TAB>script:alert(1)"} — does bypass the filter and emits <p href="java<TAB>script:alert(1)">, but href on a <p> is inert: there is nothing to navigate. (An earlier draft of this report described this as a "simpler, unconditional variant" of the attack; it is a filter bypass, not an XSS.)
  • <img src> is unaffected, since ImageRenderer unconditionally overwrites src from the core URL regardless of the safety verdict.

Recommended Fix

Normalize inside RegexHelper::isLinkPotentiallyUnsafe() before testing, mirroring the WHATWG URL parser's own normalization. This covers both variants, fixes every call site at once (LinkRenderer, ImageRenderer, and any third-party callers), and needs no changes in the Attributes extension.

Affected Versions

>= 1.5.0, <= 2.8.3 - every release that ships the AttributesExtension. Verified by installing each version and rendering the payloads with allow_unsafe_links => false. The attribute-value grammar (PARTIAL_DOUBLEQUOTEDVALUE = '"[^"]*"') has accepted raw control bytes since the extension was introduced, and none of the intervening parser rewrites narrowed it.

Prior Related Advisories

GHSA-3527-qv2q-pfvx / CVE-2025-46734 fixed a different Attributes-extension XSS (unallowlisted on* handlers, href/src not respecting allow_unsafe_links at all) in v2.7.0. This issue bypasses the specific href/src protection that fix introduced (the control-byte normalization gap was not part of that fix) - but the obfuscated inputs also work on older versions.

Show details on source website

{
  "affected": [
    {
      "database_specific": {
        "last_known_affected_version_range": "\u003c= 2.8.3"
      },
      "package": {
        "ecosystem": "Packagist",
        "name": "league/commonmark"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "1.5.0"
            },
            {
              "fixed": "2.9.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-71478"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-692",
      "CWE-79",
      "CWE-86"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-08-06T20:30:39Z",
    "nvd_published_at": null,
    "severity": "MODERATE"
  },
  "details": "\ufeff## Summary\n\nThe `AttributesExtension`\u0027s `href`/`src` unsafe-link filter (`AttributesHelper::filterAttributes()`) can be bypassed by embedding control bytes in a `javascript:` URL that browsers discard before parsing the scheme. Two variants:\n\n- **Tab/newline inside the scheme** \u2014 a literal ASCII TAB (0x09), CR (0x0D), or LF (0x0A), e.g. `java\u003cTAB\u003escript:alert(1)`. Per the WHATWG URL Standard\u0027s \"basic URL parser\" step 3, browsers \"remove all ASCII tab or newline from input\".\n- **Leading C0 controls** \u2014 e.g. `\u003c0x01\u003ejavascript:alert(1)`. Per step 1 of the same algorithm, browsers remove any leading or trailing C0 control or space. (A leading *space* alone does not bypass, because `parseAttributes()` already `trim()`s the value; other C0 bytes are not trimmed.)\n\nThe filter is a literal anchored-prefix regex (`RegexHelper::isLinkPotentiallyUnsafe()` / `REGEX_UNSAFE_PROTOCOL`) that matches neither obfuscated form, so in both cases the browser still executes `javascript:alert(1)`.\n\n**This is confirmed reproducible even with `allow_unsafe_links =\u003e false` set** \u2014 i.e. even applications that have followed the library\u0027s own documented hardening guidance for untrusted input remain exploitable.\n\nThis is a *sibling gap* in the same defense that CVE-2025-46734 (GHSA-3527-qv2q-pfvx) fixed in v2.7.0 \u2014 that fix made `href`/`src` respect `allow_unsafe_links`, but did not normalize control bytes before checking, so these obfuscation techniques were never covered.\n\n## Vulnerability\n\n**Files**:\n- `src/Util/RegexHelper.php:69` (`REGEX_UNSAFE_PROTOCOL`), `:239-242` (`isLinkPotentiallyUnsafe()`)\n- `src/Extension/Attributes/Util/AttributesHelper.php:149-179` (`filterAttributes()`)\n\n**CWE**: CWE-79 (Improper Neutralization of Input During Web Page Generation / XSS) \u2014 primary\n- CWE-692 (Incomplete Denylist to Cross-Site Scripting) \u2014 the anchored-prefix denylist in `REGEX_UNSAFE_PROTOCOL` is incomplete. This is a composite of CWE-184 and CWE-79, so it captures the full \"incomplete denylist \u2192 XSS\" chain on its own.\n- CWE-86 (Improper Neutralization of Invalid Characters in Identifiers in Web Pages) \u2014 the specific evasion technique: control bytes embedded within the URI scheme identifier, which the browser strips before resolving it.\n\n### Root Cause\n```php\n// src/Util/RegexHelper.php\npublic const REGEX_UNSAFE_PROTOCOL = \u0027/^(?:javascript|vbscript|file|data):/i\u0027;\n\npublic static function isLinkPotentiallyUnsafe(string $url): bool\n{\n    return \\preg_match(self::REGEX_UNSAFE_PROTOCOL, $url) !== 0 \u0026\u0026 \\preg_match(self::REGEX_SAFE_DATA_PROTOCOL, $url) === 0;\n}\n\n// src/Extension/Attributes/Util/AttributesHelper.php\nforeach ($attributes as $name =\u003e $value) {\n    $attrNameLower = \\strtolower($name);\n    if (! $allowUnsafeLinks \u0026\u0026 ($attrNameLower === \u0027href\u0027 || $attrNameLower === \u0027src\u0027) \u0026\u0026 \\is_string($value) \u0026\u0026 RegexHelper::isLinkPotentiallyUnsafe($value)) {\n        unset($attributes[$name]);\n        continue;\n    }\n    ...\n```\nThe Attributes extension\u0027s own quote-value grammar (`PARTIAL_DOUBLEQUOTEDVALUE = \u0027\"[^\"]*\"\u0027`) accepts any byte except `\"` inside quotes, including raw tab/CR/LF and other C0 controls, and `parseAttributes()` only `trim()`s (leading/trailing, and only the default charlist `\" \\t\\n\\r\\0\\x0B\"` \u2014 so a leading `\\x01` survives). Critically, **the core Markdown link-destination path (`LinkParserHelper` \u2192 `UrlEncoder::unescapeAndEncode()`) percent-encodes every control byte before this same safety check ever runs \u2014 but the Attributes extension\u0027s `href`/`src` handling has no equivalent normalization step**, so the raw control byte reaches both the check and the final HTML output (`Xml::escape()` only escapes `\u0026 \u003c \u003e \" \u0027`, not tab/CR/LF, since they\u0027re legal bytes inside an HTML attribute).\n\n### Attack Scenario\n1. An application enables the (commonly-used) `AttributesExtension` and sets `allow_unsafe_links =\u003e false` \u2014 the project\u0027s own documented hardening step for untrusted input.\n2. An attacker submits Markdown: `[Click me](javascript:alert(0)){href=\"java\u003cTAB\u003escript:alert(document.cookie)\"}` (TAB is one literal 0x09 byte).\n3. The library emits `\u003ca href=\"java\u003cTAB\u003escript:alert(document.cookie)\"\u003eClick me\u003c/a\u003e` \u2014 `isLinkPotentiallyUnsafe()` doesn\u0027t match the tab-split scheme, so the filter takes no action.\n4. A victim viewing/clicking the link has the browser strip the embedded TAB and execute `javascript:alert(document.cookie)` in the victim\u0027s session \u2014 stored XSS, cookie theft, account takeover potential.\n\n**Why the payload needs an unsafe core destination.** Step 2 above deliberately uses `[Click me](javascript:alert(0))` rather than a normal link. `LinkRenderer` overwrites `attrs[\u0027href\u0027]` with the node\u0027s own URL *unless* that URL is itself judged unsafe \u2014 so `[x](https://example.com){href=\"java\u003cTAB\u003escript:...\"}` renders the harmless `href=\"https://example.com\"`, and an empty destination `[x](){href=\"...\"}` renders `href=\"\"`. The attacker therefore supplies a core destination that the filter *does* catch, which suppresses the overwrite and lets the attribute-supplied `href` reach the final tag. This is no obstacle in practice \u2014 the attacker writes the entire Markdown document.\n\nTwo related forms that are **not** exploitable, noted so the fix isn\u0027t over-scoped:\n\n- Attaching the attribute to a non-link block \u2014 `hi {href=\"java\u003cTAB\u003escript:alert(1)\"}` \u2014 does bypass the filter and emits `\u003cp href=\"java\u003cTAB\u003escript:alert(1)\"\u003e`, but `href` on a `\u003cp\u003e` is inert: there is nothing to navigate. (An earlier draft of this report described this as a \"simpler, unconditional variant\" of the attack; it is a filter bypass, not an XSS.)\n- `\u003cimg src\u003e` is unaffected, since `ImageRenderer` unconditionally overwrites `src` from the core URL regardless of the safety verdict.\n\n### Recommended Fix\n\nNormalize inside `RegexHelper::isLinkPotentiallyUnsafe()` before testing, mirroring the WHATWG URL parser\u0027s own normalization. This covers both variants, fixes every call site at once (`LinkRenderer`, `ImageRenderer`, and any third-party callers), and needs no changes in the Attributes extension.\n\n## Affected Versions\n\n**`\u003e= 1.5.0, \u003c= 2.8.3`** - every release that ships the `AttributesExtension`. Verified by installing each version and rendering the payloads with `allow_unsafe_links =\u003e false`. The attribute-value grammar (`PARTIAL_DOUBLEQUOTEDVALUE = \u0027\"[^\"]*\"\u0027`) has accepted raw control bytes since the extension was introduced, and none of the intervening parser rewrites narrowed it.\n\n## Prior Related Advisories\n\nGHSA-3527-qv2q-pfvx / CVE-2025-46734 fixed a different Attributes-extension XSS (unallowlisted `on*` handlers, `href`/`src` not respecting `allow_unsafe_links` at all) in v2.7.0. This issue bypasses the specific `href`/`src` protection that fix introduced (the control-byte normalization gap was not part of that fix) - but the obfuscated inputs also work on older versions.",
  "id": "GHSA-29pj-957v-52mc",
  "modified": "2026-08-06T20:30:40Z",
  "published": "2026-08-06T20:30:39Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/thephpleague/commonmark/security/advisories/GHSA-29pj-957v-52mc"
    },
    {
      "type": "WEB",
      "url": "https://github.com/thephpleague/commonmark/commit/493a5aa7d65754b73846006eaff9c2c4431a8e2c"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/thephpleague/commonmark"
    },
    {
      "type": "WEB",
      "url": "https://github.com/thephpleague/commonmark/releases/tag/2.9.0"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N",
      "type": "CVSS_V3"
    }
  ],
  "summary": "league/commonmark: AttributesExtension href/src unsafe-link filter bypass via embedded control bytes"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Detection rules are retrieved from Rulezet.

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…