GHSA-FPF4-VWCP-V4HP

Vulnerability from github – Published: 2026-10-08 16:48 – Updated: 2026-10-08 16:48
VLAI
Summary
Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`
Details

Summary

The local web-fetch tool (web_fetch_tool, also used as the WebFetch capability's local fallback) processed responses with several steps whose running time grows quadratically with the size of certain server-controlled inputs, and ran them on the event loop: decoding the body with whichever charset the server declared, extracting the page title with a backtracking regular expression, and converting the HTML to markdown. An application that exposes this tool to untrusted prompts can be steered to fetch an attacker-controlled page of a megabyte or two that blocks the event loop for minutes, stalling every other coroutine in the process — other agent runs, other requests being served — for the duration.

This is an availability issue only. SSRF protections and the download size limit introduced in GHSA-v2xh-2vp8-57h8 are unaffected; that limit bounds how much is downloaded, not how long the response takes to process.

Details

Title extraction used a backtracking pattern over the raw response body, so a body made of repeated unterminated tag openings cost time proportional to the square of its size. The HTML-to-markdown conversion had the same shape in three of its steps: normalizing whitespace, stripping preformatted blocks, and numbering ordered lists all took time proportional to the square of a run of spaces or a list's length. All of it ran on the event loop, and the regex steps hold the interpreter lock even when moved off it, so the whole process paid for the size of a server-controlled response.

The response body was also decoded on the event loop with the codec named by the charset parameter of the response's Content-Type, looked up in Python's codec registry. That registry includes punycode, whose decoder takes time proportional to the square of its input: a response of about one megabyte labelled charset=punycode blocked the event loop for roughly half a minute, with no HTML required. The registry also includes codecs that aren't text encodings at all, such as rot_13 and base64_codec; a response labelled with one of those raised an unexpected exception out of the tool, aborting the agent run that fetched it.

Separately, the HTML-to-markdown conversion recursed once per nested element, so a page nested a few hundred elements deep raised a RecursionError out of the tool, aborting the agent run that fetched it. A JSON response nested deeper than the interpreter allows did the same. These only affect that one run.

Who Is Affected

You are affected if your application registers the local web-fetch tool (or relies on the WebFetch capability's local fallback) and exposes the agent to untrusted prompts. allowed_domains narrows the exposure to pages on those domains but does not remove it. Applications that only fetch developer-controlled URLs are not exposed to the model-chosen attack path.

Remediation

Upgrade to a patched version. The title is now found with a single linear scan, the conversion steps above run in linear time, and decoding, title extraction and conversion all run in a worker thread. A charset naming a codec that isn't a text encoding, and a page too deeply nested to convert, are reported back to the model as a failed fetch instead of aborting the run; a JSON body too deeply nested to parse is returned as plain text.

Credits

Reported privately by @BrianWillows, whose report covered the quadratic title extraction. The response decoding, the codecs that are not text encodings, and the quadratic steps in the HTML-to-markdown conversion were found while fixing it.

Show details on source website

{
  "affected": [
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "pydantic-ai"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "1.77.0"
            },
            {
              "fixed": "1.107.6"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "pydantic-ai"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "2.0.0b1"
            },
            {
              "fixed": "2.44.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "pydantic-ai-slim"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "1.77.0"
            },
            {
              "fixed": "1.107.6"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    },
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "pydantic-ai-slim"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "2.0.0b1"
            },
            {
              "fixed": "2.44.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-107290"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-1333",
      "CWE-407"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-10-08T16:48:25Z",
    "nvd_published_at": null,
    "severity": "MODERATE"
  },
  "details": "### Summary\n\nThe local web-fetch tool (`web_fetch_tool`, also used as the `WebFetch` capability\u0027s local fallback) processed responses with several steps whose running time grows quadratically with the size of certain server-controlled inputs, and ran them on the event loop: decoding the body with whichever charset the server declared, extracting the page title with a backtracking regular expression, and converting the HTML to markdown. An application that exposes this tool to untrusted prompts can be steered to fetch an attacker-controlled page of a megabyte or two that blocks the event loop for minutes, stalling every other coroutine in the process \u2014 other agent runs, other requests being served \u2014 for the duration.\n\nThis is an **availability** issue only. SSRF protections and the download size limit introduced in GHSA-v2xh-2vp8-57h8 are unaffected; that limit bounds how much is downloaded, not how long the response takes to process.\n\n### Details\n\nTitle extraction used a backtracking pattern over the raw response body, so a body made of repeated unterminated tag openings cost time proportional to the square of its size. The HTML-to-markdown conversion had the same shape in three of its steps: normalizing whitespace, stripping preformatted blocks, and numbering ordered lists all took time proportional to the square of a run of spaces or a list\u0027s length. All of it ran on the event loop, and the regex steps hold the interpreter lock even when moved off it, so the whole process paid for the size of a server-controlled response.\n\nThe response body was also decoded on the event loop with the codec named by the `charset` parameter of the response\u0027s `Content-Type`, looked up in Python\u0027s codec registry. That registry includes `punycode`, whose decoder takes time proportional to the square of its input: a response of about one megabyte labelled `charset=punycode` blocked the event loop for roughly half a minute, with no HTML required. The registry also includes codecs that aren\u0027t text encodings at all, such as `rot_13` and `base64_codec`; a response labelled with one of those raised an unexpected exception out of the tool, aborting the agent run that fetched it.\n\nSeparately, the HTML-to-markdown conversion recursed once per nested element, so a page nested a few hundred elements deep raised a `RecursionError` out of the tool, aborting the agent run that fetched it. A JSON response nested deeper than the interpreter allows did the same. These only affect that one run.\n\n### Who Is Affected\n\nYou are affected if your application registers the local web-fetch tool (or relies on the `WebFetch` capability\u0027s local fallback) and exposes the agent to untrusted prompts. `allowed_domains` narrows the exposure to pages on those domains but does not remove it. Applications that only fetch developer-controlled URLs are not exposed to the model-chosen attack path.\n\n### Remediation\n\nUpgrade to a patched version. The title is now found with a single linear scan, the conversion steps above run in linear time, and decoding, title extraction and conversion all run in a worker thread. A charset naming a codec that isn\u0027t a text encoding, and a page too deeply nested to convert, are reported back to the model as a failed fetch instead of aborting the run; a JSON body too deeply nested to parse is returned as plain text.\n\n### Credits\n\nReported privately by @BrianWillows, whose report covered the quadratic title extraction. The response decoding, the codecs that are not text encodings, and the quadratic steps in the HTML-to-markdown conversion were found while fixing it.",
  "id": "GHSA-fpf4-vwcp-v4hp",
  "modified": "2026-10-08T16:48:26Z",
  "published": "2026-10-08T16:48:25Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/security/advisories/GHSA-fpf4-vwcp-v4hp"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/pull/8397"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/pull/8399"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/pull/8418"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/pull/8433"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/pull/8434"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/commit/2faa6181d8a17d83bc9516d035c5270db8730fa0"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/commit/9cdc952e4c3319e85a3e04f2de49fbbb765bd38b"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/commit/a93ea5226be1e93ae13131ae3f22287190411389"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/commit/c3fd1cc1f15fdbf750d78e4e3ec1e8b4d6a3d920"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/commit/fb92ccfc3ca2735dab877e2ed73856681bf72ad1"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/pydantic/pydantic-ai"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/releases/tag/v1.107.6"
    },
    {
      "type": "WEB",
      "url": "https://github.com/pydantic/pydantic-ai/releases/tag/v2.44.0"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H",
      "type": "CVSS_V3"
    }
  ],
  "summary": "Pydantic AI: Event loop blocked by quadratic title extraction in `web_fetch`"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…