The text_match.py module#
Summary#
Return True iff Levenshtein distance between |
|
Normalize |
|
Return |
Description#
Shared string-matching utilities for fuzzy/typo-tolerant lookup.
These helpers are used everywhere the agent has to reconcile an LLM-supplied identifier against a finite, authoritative set of known strings — material names, allowed values of an enum-typed setting, NamedObject keys, dict-key schemas, command kwargs, etc.
Keeping the implementation in one place avoids the trap we hit
historically: _check_value learned to tolerate the LLM’s
"least-squares-cell-based" → "least-square-cell-based" typo
while copy_material could not even resolve "water-vapour" →
"water-vapor". Both are the same single-edit problem; both should
share the same solution.
- Public API:
edit_distance_le_one()— fast≤ 1Levenshtein verdict.fuzzy_normalize()— return the canonical spelling if and only if exactly one entry inallowedis within edit distance 1 (case-insensitive), elseNone.sanitize_named_object_key()— collapse whitespace in aNamedObject key to the idiomatic hyphen separator.
Module detail#
- text_match.edit_distance_le_one(a: str, b: str) bool#
Return True iff Levenshtein distance between
aandbis at most 1.Faster than computing the full distance because we only need the
≤ 1verdict.
- text_match.fuzzy_normalize(value: str, allowed: collections.abc.Iterable[str]) str | None#
Normalize
valueto a canonical spelling fromallowed.It returns the canonical spelling for
valueif exactly one entry inallowedis within edit distance 1 (case-insensitive), elseNone.Uniqueness is required so a near-spelling never silently overwrites a value when two candidates are equally plausible — surface the ambiguity to the caller instead.
- text_match.sanitize_named_object_key(name: str, *, replacement: str = '-') tuple[str, str | None]#
Return
(sanitized_name, notice_or_None)for a NamedObject key.Some backends reject whitespace in NamedObject instance keys (BC names, cell-zone names, material names, named-expression keys, …). Internal whitespace runs are collapsed to
replacement(hyphen by default — the idiomatic separator, e.g.pressure-outlet-1,phase-1,cold-inlet) and leading or trailing whitespace is stripped.The function is conservative — it ONLY rewrites whitespace. Other illegal characters (
/,\\,:, brackets, …) are rare in natural-language intent and are deliberately left alone so a deterministic rewrite never corrupts a name the user intentionally typed; the backend surfaces those at apply time.Returns a 2-tuple:
sanitized_name— the cleaned name (equal tonamewhen no rewrite was needed).notice— a one-line human-readable explanation of the rewrite, orNonewhen the input was already clean.
Non-string inputs are returned unchanged with
notice=Noneso callers can pipe values through this helper unconditionally.Example:
>>> sanitize_named_object_key("oil inlet") ('oil-inlet', "name 'oil inlet' contained whitespace; ...") >>> sanitize_named_object_key("phase-1") ('phase-1', None)