<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://vulnerability.circl.lu/rss/recent/all/10</id>
  <title>Most recent entries from all</title>
  <updated>2026-09-30T04:43:55.139344+00:00</updated>
  <author>
    <name>Vulnerability-Lookup</name>
    <email>info@circl.lu</email>
  </author>
  <link href="https://vulnerability.circl.lu" rel="alternate"/>
  <generator uri="https://lkiesow.github.io/python-feedgen" version="1.0.0">python-feedgen</generator>
  <subtitle>Contains only the most 10 recent entries.</subtitle>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-81723</id>
    <title>BREW-acronym-CVE-2026-81723 — NLTK: Quadratic CPU Exhaustion in `XMLCorpusView._read_xml_fragment()`</title>
    <updated>2026-09-30T04:43:55.144776+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> Homebrew: acronym</p>
<p>## Summary</p>
<p>`XMLCorpusView._read_xml_fragment()` reads a corpus file in 1 KiB blocks, appending
each block to a growing `fragment` string, then calls `_VALID_XML_RE.match(fragment)`
on the full accumulated buffer every iteration. Because each iteration rescans the
entire accumulated fragment, the total amount of work grows quadratically with input
size.</p>
<p>Commit `c9c332284` (CWE-1333) made each `match()` call linear. The quadratic behavior
is separate: the loop calls `match()` once per 1 KiB block, each time on a longer
buffer.</p>
<p>On the test system, an 8 MiB malformed XML file consumed approximately 48 CPU-seconds
through the public `BNCCorpusReader.words()` API with no source modification. Absolute
timings vary by hardware. `_read_xml_fragment()` imposes no limit on fragment size or
iteration count.</p>
<p>## Details</p>
<p>**File:** `nltk/corpus/reader/xmldocs.py`  
**Function:** `XMLCorpusView._read_xml_fragment()`, lines 261–308</p>
<p>The relevant loop:</p>
<p>```python
fragment = ""
while True:
    fragment += stream.read(self._BLOCK_SIZE)      # grows by 1 KiB per iteration
    if self._VALID_XML_RE.match(fragment):         # rescans full buffer each time
        return fragment
    ...
    last_open_bracket = fragment.rfind("&lt;")
    if last_open_bracket &gt; 0:                      # False for single-'&lt;' payload
        if self._VALID_XML_RE.match(fragment[:last_open_bracket]):
            return ...
    # loop continues
```</p>
<p>For a payload of `b'&lt;' + b'a' * (N-1)`:</p>
<p>- For this malformed input, `…</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-81723"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/fkie_cve-2026-81723</id>
    <title>fkie_cve-2026-81723</title>
    <updated>2026-09-30T04:43:55.144941+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml">
        <p>NLTK versions before 3.10.3 contain a quadratic CPU exhaustion vulnerability in XMLCorpusView._read_xml_fragment() that rescans accumulated XML fragments on every 1 KiB block read. Attackers can provide malformed XML corpus files to cause severe CPU consumption and denial of service through affected readers like BNCCorpusReader.</p>
      </div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/fkie_cve-2026-81723"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/ghsa-hqv3-xm29-p9hq</id>
    <title>Withdrawn: GHSA-hqv3-xm29-p9hq — Duplicate Advisory: Quadratic CPU Exhaustion in `XMLCorpusView._read_xml_fragment()`</title>
    <updated>2026-09-30T04:43:55.144999+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Withdrawn by the publisher.</strong></p>
<p><strong>Affected:</strong> PyPI: nltk</p>
<p>## Duplicate Advisory</p>
<p>This advisory has been withdrawn because it is a duplicate of GHSA-vp2x-qp44-57v7. This link is maintained to preserve external references.</p>
<p>## Original Description
NLTK versions before 3.10.3 contain a quadratic CPU exhaustion vulnerability in XMLCorpusView._read_xml_fragment() that rescans accumulated XML fragments on every 1 KiB block read. Attackers can provide malformed XML corpus files to cause severe CPU consumption and denial of service through affected readers like BNCCorpusReader.</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/ghsa-hqv3-xm29-p9hq"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/pysec-2026-3871</id>
    <title>PYSEC-2026-3871 — NLTK: Quadratic CPU Exhaustion in `XMLCorpusView._read_xml_fragment()`</title>
    <updated>2026-09-30T04:43:55.145049+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> PyPI: nltk</p>
<p>## Summary</p>
<p>`XMLCorpusView._read_xml_fragment()` reads a corpus file in 1 KiB blocks, appending
each block to a growing `fragment` string, then calls `_VALID_XML_RE.match(fragment)`
on the full accumulated buffer every iteration. Because each iteration rescans the
entire accumulated fragment, the total amount of work grows quadratically with input
size.</p>
<p>Commit `c9c332284` (CWE-1333) made each `match()` call linear. The quadratic behavior
is separate: the loop calls `match()` once per 1 KiB block, each time on a longer
buffer.</p>
<p>On the test system, an 8 MiB malformed XML file consumed approximately 48 CPU-seconds
through the public `BNCCorpusReader.words()` API with no source modification. Absolute
timings vary by hardware. `_read_xml_fragment()` imposes no limit on fragment size or
iteration count.</p>
<p>## Details</p>
<p>**File:** `nltk/corpus/reader/xmldocs.py`  
**Function:** `XMLCorpusView._read_xml_fragment()`, lines 261–308</p>
<p>The relevant loop:</p>
<p>```python
fragment = ""
while True:
    fragment += stream.read(self._BLOCK_SIZE)      # grows by 1 KiB per iteration
    if self._VALID_XML_RE.match(fragment):         # rescans full buffer each time
        return fragment
    ...
    last_open_bracket = fragment.rfind("&lt;")
    if last_open_bracket &gt; 0:                      # False for single-'&lt;' payload
        if self._VALID_XML_RE.match(fragment[:last_open_bracket]):
            return ...
    # loop continues
```</p>
<p>For a payload of `b'&lt;' + b'a' * (N-1)`:</p>
<p>- For this malformed input, `…</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/pysec-2026-3871"/>
  </entry>
  <entry>
    <id>https://vulnerability.circl.lu/vuln/ubuntu-cve-2026-81723</id>
    <title>UBUNTU-CVE-2026-81723</title>
    <updated>2026-09-30T04:43:55.145143+00:00</updated>
    <content type="xhtml">
      <div xmlns="http://www.w3.org/1999/xhtml"><p><strong>Affected:</strong> Ubuntu:Pro:14.04:LTS: nltk, Ubuntu:Pro:16.04:LTS: nltk, Ubuntu:Pro:18.04:LTS: nltk, Ubuntu:Pro:20.04:LTS: nltk, Ubuntu:Pro:22.04:LTS: nltk, Ubuntu:Pro:24.04:LTS: nltk, Ubuntu:Pro:26.04:LTS: nltk</p>
<p>NLTK versions before 3.10.3 contain a quadratic CPU exhaustion vulnerability in XMLCorpusView._read_xml_fragment() that rescans accumulated XML fragments on every 1 KiB block read. Attackers can provide malformed XML corpus files to cause severe CPU consumption and denial of service through affected readers like BNCCorpusReader.</p></div>
    </content>
    <link href="https://vulnerability.circl.lu/vuln/ubuntu-cve-2026-81723"/>
  </entry>
</feed>
