<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from all</title>
    <link>https://vulnerability.circl.lu</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Tue, 06 Oct 2026 01:25:31 +0000</lastBuildDate>
    <item>
      <title>BREW-acronym-CVE-2026-12061 — Natural Language Toolkit (NLTK): ReDoS in NLTK ReviewsCorpusReader FEATURES regex</title>
      <link>https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-12061</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; Homebrew: acronym&lt;/p&gt;
&lt;p&gt;### Summary
`ReviewsCorpusReader` extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then `[+2]`) from each review line, using the module-level `FEATURES` regex. The feature-label sub-pattern is unbounded — an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal `[`. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs `reviews()`, `features()`, and `sents()`.&lt;/p&gt;
&lt;p&gt;### Details
The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal `[`. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. `re.findall` repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.&lt;/p&gt;
&lt;p&gt;### PoC
```
import multiprocessing as mp
import re
import time&lt;/p&gt;
&lt;p&gt;# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r&amp;#34;((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]&amp;#34;)&lt;/p&gt;
&lt;p&gt;#…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; Homebrew: acronym&lt;/p&gt;
&lt;p&gt;### Summary
`ReviewsCorpusReader` extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then `[+2]`) from each review line, using the module-level `FEATURES` regex. The feature-label sub-pattern is unbounded — an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal `[`. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs `reviews()`, `features()`, and `sents()`.&lt;/p&gt;
&lt;p&gt;### Details
The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal `[`. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. `re.findall` repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.&lt;/p&gt;
&lt;p&gt;### PoC
```
import multiprocessing as mp
import re
import time&lt;/p&gt;
&lt;p&gt;# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r&amp;#34;((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]&amp;#34;)&lt;/p&gt;
&lt;p&gt;#…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/brew-acronym-cve-2026-12061</guid>
    </item>
    <item>
      <title>PYSEC-2026-3582 — Natural Language Toolkit (NLTK): ReDoS in NLTK ReviewsCorpusReader FEATURES regex</title>
      <link>https://vulnerability.circl.lu/vuln/pysec-2026-3582</link>
      <description>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: nltk&lt;/p&gt;
&lt;p&gt;### Summary
`ReviewsCorpusReader` extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then `[+2]`) from each review line, using the module-level `FEATURES` regex. The feature-label sub-pattern is unbounded — an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal `[`. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs `reviews()`, `features()`, and `sents()`.&lt;/p&gt;
&lt;p&gt;### Details
The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal `[`. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. `re.findall` repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.&lt;/p&gt;
&lt;p&gt;### PoC
```
import multiprocessing as mp
import re
import time&lt;/p&gt;
&lt;p&gt;# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r&amp;#34;((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]&amp;#34;)&lt;/p&gt;
&lt;p&gt;#…&lt;/p&gt;</description>
      <content:encoded>&lt;p&gt;&lt;strong&gt;Affected:&lt;/strong&gt; PyPI: nltk&lt;/p&gt;
&lt;p&gt;### Summary
`ReviewsCorpusReader` extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then `[+2]`) from each review line, using the module-level `FEATURES` regex. The feature-label sub-pattern is unbounded — an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal `[`. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs `reviews()`, `features()`, and `sents()`.&lt;/p&gt;
&lt;p&gt;### Details
The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal `[`. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. `re.findall` repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.&lt;/p&gt;
&lt;p&gt;### PoC
```
import multiprocessing as mp
import re
import time&lt;/p&gt;
&lt;p&gt;# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r&amp;#34;((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]&amp;#34;)&lt;/p&gt;
&lt;p&gt;#…&lt;/p&gt;</content:encoded>
      <guid isPermaLink="false">https://vulnerability.circl.lu/vuln/pysec-2026-3582</guid>
    </item>
  </channel>
</rss>
