Download the PHP package loupe/matcher without Composer
On this page you can find all versions of the php package loupe/matcher. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download loupe/matcher
More information about loupe/matcher
Files in loupe/matcher
Package matcher
Short Description Tokenize, decompose, highlight and crop around text and search terms
License MIT
Informations about the package matcher
Loupe Matcher
Loupe Matcher turns plain search queries and arbitrary text into precise, user-friendly matches: tokenize phrases and negations, normalize language-specific spelling variants, decompose compound words, calculate match spans, and format the result with highlighting, cropping, and truncation.
Installation
Quick Start
Here's a simple example of how to use Loupe Matcher to highlight search terms in a text document and crop around the highlights:
Core Components
Tokenizer
Purpose: Breaks text into searchable tokens (words, phrases, terms) for accurate matching.
The Tokenizer converts strings into TokenCollection objects, handling:
- Word boundaries using
ext-intlrules - Phrase groups (quoted terms like
"exact phrase") - Negated terms (prefixed with
-) - Locale-specific tokenization
- Locale-specific term decomposition
If you want to configure the way the Tokenizer handles locale specifics (such as decomposition or normalization), you
can provide your own implementation of the LocaleConfigurationInterface or use any of the pre-built configurations shipped
with this library. There are currently the following:
- English: Handles decomposition (
toothbrush->tooth,brush) - German: Handles normalization of German umlauts as well as
ßand also decomposition (Zeitungspapier->zeitung,papier)
Checkout the separate docs on decomposition if you want to improve the existing locale configurations or add support for a new one!
FastSet dictionary cache
The packaged dictionaries are compressed. On their first use, FastSet creates optimized index files. By default those
files are written next to the packaged dictionary, which is usually inside vendor/. Pass a writable cache directory
when creating a preconfigured locale to store them elsewhere:
The locale is appended to the supplied directory, so German files in this example are stored in
/var/cache/my-app/loupe-matcher/de. Locale-specific directories are created automatically. You can also pass the same
value as the $fastSetCacheDirectory second argument to new English() or new German().
Matcher
Purpose: Finds which tokens in your text match the search query.
The Matcher compares tokenized text against search terms, with support for:
- Stop word filtering (ignore common words like "the", "and")
- Match span calculation (start/end positions)
- Flexible matching between token collections
Formatter
Purpose: Combines matching and highlighting to create formatted output with context.
The Formatter orchestrates the entire process:
- Highlights matched terms with HTML tags
- Crops text to show relevant context around matches
- Truncates long text while preserving word boundaries and highlights
- Configurable through
FormatterOptions
Match Prioritization
By default, cropping emits snippets around every match cluster and truncation cuts from the start. Enabling withEnableMatchPrioritization() will attempt to choose the most relevant window(s) for display. Windows are scored by distinct query terms hit, then total matches, then density.
- Cropping now finds the best windows around matches, limits each window to
crop_lengthand shows up tocrop_max_fragmentswindows in document order. - Truncation picks a single window centered on the best cluster of matches and falls back to truncating from the start if no matches are found in the attribute.
Advanced Usage
Custom Tokenizer
Implement TokenizerInterface for specialized tokenization:
Pre-highlighted Text Cropping
When you already have highlighted text that needs cropping:
Using Pre-calculated Matches
When you already have a TokenCollection of matches (e.g., from a previous search operation or external source), you can format text directly without re-calculating matches. This approach is useful when your search engine already provides match information or you want to cache match results for performance.