Skip to main content

node-re2 CVE-2026-71498

| EUVDEUVD-2026-54165 MEDIUM
Out-of-bounds Read (CWE-125)
2026-08-06 https://github.com/uhop/node-re2 GHSA-j4r3-hg7j-8chg
5.1
CVSS 3.1 · Vendor: https://github.com/uhop/node-re2
Share

Severity by source

Vendor (https://github.com/uhop/node-re2) PRIMARY
5.1 MEDIUM
AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L
vuln.today AI
5.3 MEDIUM

Network vector reflects typical web-service deployment; no auth or grooming required; impact is heap disclosure only, no integrity or reliable availability impact.

3.1 AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N
4.0 AV:N/AC:L/AT:N/PR:N/UI:N/VC:L/VI:N/VA:N/SC:N/SI:N/SA:N
Red Hat
5.1 MEDIUM
qualitative

Primary rating from Vendor (https://github.com/uhop/node-re2).

CVSS VectorVendor: https://github.com/uhop/node-re2

Attack Vector
Local
Attack Complexity
Low
Privileges Required
None
User Interaction
None
Scope
Unchanged
Confidentiality
Low
Integrity
None
Availability
Low

Lifecycle Timeline

2
Source Code Evidence Fetched
Aug 06, 2026 - 22:01 vuln.today
Analysis Generated
Aug 06, 2026 - 22:01 vuln.today

Blast Radius

ecosystem impact
† from your stack dependencies † transitive graph · vuln.today resolves 4-path depth
  • 534 npm packages depend on re2 (114 direct, 428 indirect)

Ecosystem-wide dependent count for version 1.26.1.

DescriptionCVE.org

Summary

re2 infers a character's byte length from its UTF-8 lead byte alone, with no bound on the bytes actually remaining in the input. Buffer arguments reach the native layer verbatim - only strings are re-encoded into well-formed UTF-8 - so a Buffer whose last byte is a multi-byte lead promises continuation bytes that are not there, and the result builders read up to 3 bytes past the end of the buffer. In replace() and split() those bytes are copied into the returned Buffer, disclosing adjacent heap memory to JavaScript. The trigger is deterministic and requires no special heap grooming.

Only Buffer input is affected. String input was never at risk: re-encoding guarantees every multi-byte sequence is complete.

Root cause

getUtf8CharSize maps a lead byte to a length of 1-4 and never sees the input size:

cpp
// lib/wrapped_re2.h
inline size_t getUtf8CharSize(char ch)
{
      return ((0xE5000000 >> ((ch >> 3) & 0x1E)) & 3) + 1;
}

Callers then read that many bytes. In the zero-width branch of replace(), the guard proves only that at least *one* byte remains:

cpp
// lib/replace.cc
else if ((size_t)offset < size)
{
      auto sym_size = getUtf8CharSize(data[offset]);   // may claim up to 4 bytes
      result.append(data + offset, sym_size);          // reads data[offset .. offset + 3]
      byteIndex = offset + sym_size;
}

offset < size permits offset == size - 1, so a lead byte of 0xF0 makes append read data[size], data[size + 1] and data[size + 2].

Seven read sites shared the defect:

SiteArgumentDisclosed to JS
lib/replace.cc (zero-width branch)subjectyes
lib/replace.cc (callback replacer)subjectyes
lib/replace.cc (replacement scan)replacementyes
lib/split.ccsubjectyes
lib/pattern.cc translateRegExp (x2)patternno
lib/pattern.cc escapeRegExppatternno

Three further callers were not vulnerable, because they use the result only to advance an index and never dereference past the end: getUtf16PositionByCounter in lib/wrapped_re2.h (clamps its return to the buffer size), lib/match.cc (the value feeds RE2::Match, which rejects startpos > endpos), and the getMaxSubmatch scan in lib/replace.cc (an overshoot just ends the loop).

Proof of concept

Each call returns more bytes than were supplied; the trailing bytes are heap contents and vary between runs.

js
const RE2 = require('re2');
const hex = buf => [...buf].map(b => b.toString(16).padStart(2, '0')).join(' ');

// subject: 2 bytes in, 5 bytes out
console.log(hex(new RE2('', 'g').replace(Buffer.from([0x41, 0xf0]), '')));
// 41 f0 61 7b eb   <- last 3 bytes are adjacent heap memory

// replacement argument
console.log(hex(new RE2('A', 'g').replace(Buffer.from('A'), Buffer.from([0x42, 0xf0]))));
// 42 f0 41 26 d6

// split
console.log(new RE2('', 'g').split(Buffer.from([0x41, 0xf0])).map(hex));
// [ '41', 'f0 e2 e4 df' ]

0xC2 (2-byte lead) and 0xE2 (3-byte lead) over-read 1 and 2 bytes respectively; 0xF0 over-reads 3.

For the pattern path the over-read occurs in translateRegExp / escapeRegExp, which run before RE2 validates the pattern, but RE2 then rejects the malformed input, so the bytes are discarded rather than returned:

js
new RE2(Buffer.from([0xf0]));   // SyntaxError: invalid UTF-8 - read already happened

Impact

Information disclosure (replace, split). Up to 3 bytes of heap memory adjacent to the input buffer are returned to JavaScript per call. The read is repeatable, so an attacker who controls Buffer input and observes output can sample heap memory incrementally. What lands there depends on allocator layout and is not directly steerable, but it may include fragments of other buffers.

Out-of-bounds read (pattern compilation). No disclosure path, since the malformed pattern is rejected - but the read is still undefined behavior and can fault if the buffer ends on a page boundary.

Applications that pass only strings, or only well-formed UTF-8 buffers, are unaffected. The exposure matters most where re2 is used as intended: running patterns or subjects derived from untrusted input.

Suggested fix

Clamp the inferred character size to the bytes that actually remain, at every site whose result indexes the buffer:

cpp
inline size_t getUtf8CharSize(char ch, size_t remaining)
{
      size_t size = getUtf8CharSize(ch);
      return size < remaining ? size : remaining;
}

This is O(1) and changes no algorithm's complexity. A truncated tail then round-trips as the bytes it really holds, which preserves the documented contract that Buffer input is passed through verbatim. Rejecting malformed UTF-8 in Buffer input would also close the hole, but is a breaking API change.

Resolution

Fixed in re2@1.26.1.

All seven read sites now clamp the character size to the remaining input, so a Buffer ending in a truncated multi-byte character round-trips as its own bytes instead of reading past the end. Regression tests cover the subject, replacement and pattern positions for 2-, 3- and 4-byte leads, including partially truncated sequences.

Remediation: upgrade to re2@1.26.1 or later.

Workaround (if you cannot upgrade): pass strings rather than Buffers, or validate that Buffer input is well-formed UTF-8 before calling replace, split, or the RE2 constructor - for example Buffer.compare(Buffer.from(buf.toString('utf8')), buf) === 0.

Reported by @OvOhao in #272.

AnalysisAI

Out-of-bounds heap read in the node-re2 npm package (versions <= 1.26.0) exposes up to 3 bytes of adjacent heap memory per call to JavaScript when Buffer arguments ending in a truncated multi-byte UTF-8 lead byte are passed to replace() or split(). The overread is deterministic, requires no heap grooming, and a working proof-of-concept is included in the GitHub advisory. Applications that pass only strings - not Buffers - are entirely unaffected, but those that process untrusted Buffer input through re2 are at real risk of incremental heap disclosure. No CISA KEV listing was identified; this is a library-layer flaw fixed in re2@1.26.1.

Technical ContextAI

node-re2 (pkg:npm/re2) is a Node.js native addon wrapping Google's RE2 regex engine, used in Node.js applications as a safe, linear-time alternative to the built-in RegExp. The root cause (CWE-125: Out-of-bounds Read) is in the C++ helper getUtf8CharSize(char ch) in lib/wrapped_re2.h, which derives a character's byte length from its UTF-8 lead byte alone using a bitmask table, with no awareness of how many bytes remain in the input buffer. Node.js Buffer objects bypass the string re-encoding path and land in the native layer verbatim, so a Buffer whose last byte is a multi-byte lead byte (0xC2=2 bytes, 0xE2=3 bytes, 0xF0=4 bytes) causes callers to read and append 1-3 bytes beyond the allocated buffer. Four of seven affected read sites in replace.cc and split.cc copy the over-read bytes into the returned JavaScript Buffer, while three sites in pattern.cc trigger the read during translateRegExp/escapeRegExp but discard the result when RE2 rejects the malformed pattern. The fix in commit 9d72042 clamps inferred character size to the bytes actually remaining at every affected site.

RemediationAI

Upgrade to re2@1.26.1 or later (npm install re2@1.26.1), which clamps getUtf8CharSize() to the remaining buffer bytes at all seven affected read sites. The fix is confirmed by commit 9d72042 and the GitHub advisory GHSA-j4r3-hg7j-8chg. If an immediate upgrade is not possible, the advisory-documented workaround is to pass JavaScript strings rather than Buffer objects to re2 APIs, since strings are re-encoded into well-formed UTF-8 before reaching the native layer. Alternatively, validate that any Buffer is well-formed UTF-8 before calling replace(), split(), or the RE2 constructor using Buffer.compare(Buffer.from(buf.toString('utf8')), buf) === 0 as a guard; reject any Buffer where this check fails. The validation approach adds a full O(n) string round-trip per call, so prefer the upgrade path in production.

Vendor StatusVendor

Share

CVE-2026-71498 vulnerability details – vuln.today

This site uses cookies essential for authentication and security. No tracking or analytics cookies are used. Privacy Policy