Download the PHP package iliaal/phonetic without Composer

On this page you can find all versions of the php package iliaal/phonetic. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.

FAQ

After the download, you have to make one include require_once('vendor/autoload.php');. After that you have to import the classes with use statements.

Example:
If you use only one package a project is not needed. But if you use more then one package, without a project it is not possible to import the classes with use statements.

In general, it is recommended to use always a project to download your libraries. In an application normally there is more than one library needed.
Some PHP packages are not free to download and because of that hosted in private repositories. In this case some credentials are needed to access such packages. Please use the auth.json textarea to insert credentials, if a package is coming from a private repository. You can look here for more information.

  • Some hosting areas are not accessible by a terminal or SSH. Then it is not possible to use Composer.
  • To use Composer is sometimes complicated. Especially for beginners.
  • Composer needs much resources. Sometimes they are not available on a simple webspace.
  • If you are using private repositories you don't need to share your credentials. You can set up everything on our site and then you provide a simple download link to your team member.
  • Simplify your Composer build process. Use our own command line tool to download the vendor folder as binary. This makes your build process faster and you don't need to expose your credentials for private repositories.
Please rate this library. Is it a good library?

Informations about the package phonetic

phonetic

Native phonetic matching for PHP: Double Metaphone, Beider-Morse Phonetic Matching (BMPM), Daitch-Mokotoff Soundex, NYSIIS, and Match Rating Approach, the phonetic name-matching encoders that PHP core does not ship. It also ships comparison helpers that answer "do these two names sound alike?" directly.

PHP core has soundex() and metaphone(), but not these, which are the standard tools for fuzzy name matching, record linkage, and genealogy search across spelling and transliteration variants.

Quick Start

Install via PIE (requires PHP 8.1 or later):

Then ask whether two names sound alike, no userland matching logic required:

Choosing an algorithm

Double Metaphone BMPM Daitch-Mokotoff Soundex NYSIIS Match Rating
Output primary + alternate key language-aware token set distinct 6-digit codes single key compact codex
Two names match when keys are equal token sets intersect code sets intersect keys are equal clear the MRA similarity threshold
Strongest for English and general Latin-script names cross-language and transliteration variants (Slavic, Germanic, Hebrew, Romance) Eastern-European and Ashkenazi surnames, genealogy American/English surnames English names; ships its own similarity test
Spelling-variant recall good highest high, within its language model good good
Ambiguity handling up to 2 keys many tokens multiple codes single key single codex
Relative speed fast (1.0x) slowest (~60x) middle (~2.3x) fast (0.42x) fastest (0.24x)
Data source clean-room published algorithm Apache Commons Codec rule data Apache Commons Codec rule data clean-room published algorithm clean-room published algorithm

Rule of thumb: reach for Double Metaphone as a fast general-purpose default, BMPM when names cross languages or scripts, and Daitch-Mokotoff for Eastern-European and Jewish genealogy where it is the field standard. NYSIIS and Match Rating Approach are lighter, single-key English/American encoders, useful as alternate index keys or a second opinion alongside Double Metaphone.

API

Double Metaphone

Primary + alternate phonetic keys (Lawrence Philips). Clean-room implementation.

alternate equals primary when the algorithm produced no alternate branch. max_length caps each key (default 4; 0 or negative = unlimited).

Beider-Morse Phonetic Matching

Language-aware token set. | separates alternatives. With the default concatenation mode, ordinary words form one encoded sequence ("John Smith" becomes "ionzmit"). Recognized generic prefixes use (remainder)-(combined) groups. Matches Apache Commons Codec's default BeiderMorseEncoder.

Empty $language auto-detects; pass an exact lowercase language token for that name type (e.g. "russian", "english") to force it. Tokens are name-type-specific (GENERIC has the largest set; ASHKENAZI/SEPHARDIC are subsets). The label "any" is not forceable — see Notes.

Constants (numeric values):

Constant Value Role
BMPM_GENERIC 0 name type
BMPM_ASHKENAZI 1 name type
BMPM_SEPHARDIC 2 name type
BMPM_APPROX 10 accuracy
BMPM_EXACT 20 accuracy

Accuracy values are deliberately disjoint from name-type values so a misplaced constant such as bmpm($s, BMPM_APPROX) is rejected instead of silently selecting a name type. Prefer the constant names over hard-coded integers.

Invalid $name_type, $accuracy, or unknown $language for the name type raise ValueError. Inputs longer than 4096 bytes also raise ValueError (security bound shared with bmpm_match()).

A forced language also applies to the split variants of prefixed names (van Smith, d'Angelo). Commons Codec re-detects the language inside its prefix branch, silently ignoring the forced set there; this extension deliberately diverges and keeps it forced.

Daitch-Mokotoff Soundex

List of distinct 6-digit codes (the algorithm branches on ambiguous letters). Matches Apache Commons Codec's DaitchMokotoffSoundex in branching mode.

Empty string is the one encode-level departure from Commons Codec: this API returns []. Non-empty input that matches no rule still returns ["000000"] for encoder parity with the oracle.

Indexing caveat: the same "000000" string is also the finished code for pure-vowel inputs that did match a rule (e.g. "A"). Encode equality of "000000" does not imply a match: dm_soundex_match("A", "1") is false while both encode to ["000000"]. When building an inverted index from dm_soundex(), skip the "000000" key (or use dm_soundex_match() at query time).

NYSIIS

Single phonetic key (New York State Identification and Intelligence System), tuned for American/English surnames. Reimplementation of the published algorithm; matches Apache Commons Codec's Nysiis.

The classic algorithm truncates to 6 characters; max_length = 0 (or negative) returns the full key.

Match Rating Approach

Compact codex (Western Airlines, 1977). Pair it with its own similarity test instead of comparing codexes for equality.

Use match_rating_compare() (below) for the actual homophone decision. It applies the algorithm's length-and-rating rules that plain codex equality skips.

Comparison helpers

Each encoder produces a different output shape, so "do these sound alike?" needs the right comparison per algorithm. These helpers encapsulate that, so you don't reimplement the set-intersection or match-strength logic in userland.

Empty / unencodable never match. Across all helpers, an empty encoding or an input that produces no usable code is not a homophone of anything (including another empty/unencodable input):

Usage

For a one-off "do these sound alike?" check, use the comparison helpers directly. Each applies the correct per-algorithm logic:

For indexed lookup, encode once and store the key(s) with each record, then query by encoded value instead of re-encoding at search time. Double Metaphone gives one or two keys per name; Daitch-Mokotoff and BMPM give a set, so index every code. BMPM separates alternatives with |; ordinary words are concatenated, while recognized generic prefixes produce (remainder)-(combined) groups:

Performance

Single-name encode, warm, -O2 non-ASan PHP 8.4 on one core, over a representative mix of 18 names (best of 5 trials). Absolute time scales with input length; the relative ordering is the stable part.

encoder per call throughput relative
match_rating() ~0.043 µs ~23M/sec 0.24x
nysiis() ~0.074 µs ~13M/sec 0.42x
double_metaphone() ~0.18 µs ~5.5M/sec 1.0x
dm_soundex() ~0.41 µs ~2.4M/sec ~2.3x slower
bmpm() ~11 µs ~91k/sec ~60x slower

Match Rating and NYSIIS are short single-key passes, so they're the cheapest. Double Metaphone is a single linear pass with a primary/alternate split. Daitch-Mokotoff branches on ambiguous letters and dedups the resulting codes; a first-byte rule index keeps it fast. BMPM is the heaviest: language detection, a main transliteration pass, and two final rule passes, expanding a Cartesian product of phoneme alternatives capped at 20 per word. When you know the language, passing an explicit $language skips auto-detection and can cut bmpm time several-fold, though the gain depends on the chosen language's ruleset. Choose BMPM for recall, not throughput.

The comparison helpers cost roughly two encodes plus a cheap compare:

helper per call throughput
match_rating_compare() ~0.11 µs ~9M/sec
nysiis_match() ~0.14 µs ~7M/sec
double_metaphone_match() ~0.26 µs ~3.8M/sec
dm_soundex_match() ~0.80 µs ~1.3M/sec
bmpm_match() ~22 µs ~45k/sec

For repeated lookups against a fixed corpus, encode once and index the keys (see Usage) rather than calling a helper per candidate pair.

Notes & limitations

Input-length policy by function (the cap is per-argument, so both operands of a match/compare helper are checked):

Function(s) Max input Over the limit
bmpm, bmpm_match 4096 bytes throws ValueError
dm_soundex, dm_soundex_match 4096 bytes throws ValueError
double_metaphone, double_metaphone_match not capped encodes (bounded by memory)
nysiis, nysiis_match not capped encodes (bounded by memory)
match_rating, match_rating_compare not capped encodes (bounded by memory)

The uncapped encoders run in linear time and space, so bound untrusted input at the application layer if you feed them arbitrary-length strings.

🔗 Native PHP extensions

Companion native PHP extensions:

License

BSD 3-Clause (see LICENSE).

The Beider-Morse and Daitch-Mokotoff rule data is vendored from Apache Commons Codec under the Apache License 2.0. The complete terms are in LICENSE-APACHE; the Commons Codec attribution is in NOTICE, and Section 2 of LICENSE maps the rule data to those files. Double Metaphone, NYSIIS, and Match Rating Approach are independent implementations of their published algorithms, validated against (and edge-case-aligned with) Apache Commons Codec as the parity-test oracle (no third-party code or data).


Follow @iliaa on X • Blog • If this matched the names exact comparison missed, ⭐ star it!


All versions of phonetic with dependencies

PHP Build Version
Package Version
Requires php Version >=8.1
Composer command for our command line client (download client) This client runs in each environment. You don't need a specific PHP version etc. The first 20 API calls are free. Standard composer command

The package iliaal/phonetic contains the following files

Loading the files please wait ...