Quickstart¶
Installation¶
uv add replus
# or
pip install replus
Requires Python 3.11+. The only runtime dependency is regex.
Define pattern templates¶
The engine loads pattern templates from *.json files in a directory (or from a
plain dict of dicts). Every key defines a named fragment; fragments compose through
{{placeholders}}. Only the patterns under the special $PATTERNS key are matched
against at runtime.
patterns/date.json:
{
"day": ["3[01]", "[12][0-9]", "0?[1-9]"],
"month": ["0?[1-9]", "1[012]"],
"year": ["\\d{4}"],
"date": ["{{day}}/{{month}}/{{year}}", "{{year}}-{{month}}-{{day}}"],
"$PATTERNS": ["{{date}}"]
}
Note
Backslashes must be double-escaped in JSON: write \\d to get \d.
This compiles to a single pattern with automatically numbered named groups:
(?P<date_0>(?P<day_0>3[01]|[12][0-9]|0?[1-9])/(?P<month_0>0?[1-9]|1[012])/(?P<year_0>\d{4})|(?P<year_1>\d{4})-(?P<month_1>0?[1-9]|1[012])-(?P<day_1>3[01]|[12][0-9]|0?[1-9]))
Match and query¶
from replus import Replus
engine = Replus("patterns") # or a dict of dicts
for match in engine.parse("Look at this date: 1970-12-25"):
print(repr(match))
# <[Match date] span(19, 29): 1970-12-25>
date = match.group("date")
print(repr(date))
# <Group date_0 span(19, 29) @0: '1970-12-25'>
print(repr(date.group("day")))
# <Group day_1 span(27, 29) @0: '25'>
print(repr(date.group("month")))
# <Group month_1 span(24, 26) @0: '12'>
print(repr(date.group("year")))
# <Group year_1 span(19, 23) @0: '1970'>
You query groups by their template key (day), not the generated group name
(day_1) — the engine resolves which alternative actually matched.
Every match serializes to a plain dict or JSON:
match.serialize() # nested dict of values, offsets, and groups
match.json(indent=2)
Filtering by type¶
The stem of each template file (or the dict key) becomes the match type, so you can select which pattern families run:
engine.parse(text, filters=["date"]) # only date patterns
engine.parse(text, exclude=["cities"]) # everything but cities
The three matching methods¶
Method |
Returns |
Overlaps |
|---|---|---|
|
|
purged (longest wins) unless |
|
first |
same as |
|
lazy |
raw, unpurged |
All three accept the same keyword arguments and forward the match-time extras
(pos, endpos, overlapped, partial, concurrent, timeout) to
regex.
Note
Regex flags are a compile-time setting: pass them once when constructing the
engine (Replus("patterns", flags=regex.IGNORECASE)). They cannot be given per call,
because the patterns are already compiled by then.
Replacing inside matches¶
sub() rewrites the text captured by specific template keys, leaving everything
else — including unmatched text — untouched. This is built for normalization jobs
like OCR-error correction, where a 1 misread for an l must only be fixed
inside a matched serial number:
engine = Replus({
"serial": {
"prefix": ["[1lI]"],
"digits": ["\\d{6}"],
"serial": ["{{prefix}}{{digits}}"],
"$PATTERNS": ["{{serial}}"],
}
})
engine.sub("ID 1809900 and 1809900", {"prefix": "l"})
# 'ID l809900 and l809900'
A replacement can also be a callable receiving the Group,
which is the right tool when the fix depends on what was captured:
engine.sub(
"ID 1809900 and I809900",
{"prefix": lambda g: g.value.replace("1", "l").replace("I", "l")},
)
# 'ID l809900 and l809900'
Rules of the road:
String replacements are inserted literally — no
\g<1>-style escape processing.Every repetition of a group is replaced; a callable can branch on
g.rep_indexto target specific ones.A key defined by no pattern raises
NoSuchGroupError(typo guard); a key that simply didn’t capture in a match is skipped.Requesting overlapping targets (a group and one of its children) raises
OverlappingReplacementErrorinstead of guessing which edit wins.sub()accepts the same keyword arguments asparse()(e.g.filters), exceptoverlapped.
Continue with the template syntax reference for backreferences, unnamed groups, lookarounds, and repeated captures.