Download the PHP package daikazu/gibberish-name-detector without Composer
On this page you can find all versions of the php package daikazu/gibberish-name-detector. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download daikazu/gibberish-name-detector
More information about daikazu/gibberish-name-detector
Files in daikazu/gibberish-name-detector
Package gibberish-name-detector
Short Description Catch gibberish names like "asdfgh" in Laravel validation — flags keyboard-mash and random-string input in name fields.
License MIT
Homepage https://github.com/daikazu/gibberish-name-detector
Informations about the package gibberish-name-detector
Gibberish Name Detector
A pure-PHP gibberish / random-string detector for Laravel name fields. It returns a
calibrated probability that a short string is mashed or random input (asdfgh,
qwerty, xkjqwz) rather than a real name, using a logistic-regression classifier over
engineered features (n-gram scores + structural signals), with an optional Bloom-filter
name gazetteer that whitelists known real names.
The output is a probability in [0, 1], not a hard truth. Use it to fail validation,
soft-warn ("Did you type your name correctly?"), or flag a submission for review.
How it works
Inference is pure PHP — n-gram table lookups, a dot product through a sigmoid, and a
Bloom membership test. There is *no ML runtime, no FFI, no `ext-requirement, and no network call** on the inference path. The trained model and gazetteer ship as committed PHP array files underresources/data/(opcache-friendly; neverjson_decode`d).
analyze() runs a short-circuit ladder:
- Normalize — lowercase;
letters=[a-z];tokens= split on non-letters (soMary-Jane,O'Brien,Van Bergscore per token). - Length gate — too short → not gibberish (
reason: too_short). - Gazetteer — every token is a known name → not gibberish (
reason: gazetteer). - Fast pre-reject — extreme junk (long repeats, keyboard walks, no vowels) →
gibberish (
reason: heuristic). - Classifier — logistic regression → probability;
gibberish = p >= threshold(reason: model).
Requirements
- PHP 8.2+ (8.3+ for Laravel 13)
- Laravel 12 or 13 (the core detector also runs framework-free via plain
new)
Installation
No datasets to download. The trained model and gazetteer ship committed with the package under
resources/data/. The external name lists referenced in Regenerating the artifacts are only needed if you want to retrain from scratch.
The service provider and Gibberish facade are auto-discovered. Publish the config to
tune the threshold or messaging:
Usage
Facade
Validation
Rule object:
String rule:
Both resolve the configured singleton when given no overrides.
Configuration
config/gibberish.php:
| Key | Env | Default | Meaning |
|---|---|---|---|
threshold |
GIBBERISH_THRESHOLD |
null |
Probability cutoff. null → the model's calibrated threshold. |
min_length |
GIBBERISH_MIN_LENGTH |
3 |
Below this letter count, never flagged. |
gazetteer |
GIBBERISH_GAZETTEER |
true |
Enable the known-name whitelist. |
message |
— | … | Validation failure message (:attribute supported). |
Threshold note. The shipped model carries its own calibrated threshold (~0.46, chosen by Youden's J). Leaving
thresholdnulluses it — this is what keeps recall ≥ 90% out of the box. SetGIBBERISH_THRESHOLDto a fixed value to override (e.g. raise it for fewer false positives when hard-failing validation).
The gibberish:check command
Inspect a single value (debug / manual QA):
Caveats
- It is a probability, not the truth. Real but unusual names will sometimes score high; treat a positive as a signal (warn / review), not a certainty — especially when hard-failing validation.
- Non-Anglo names. Features are computed on
[a-z]only (accented characters are dropped), so the classifier alone can flag names likeOkonkwo. The gazetteer is the fix: it whitelists ~92k known names across many locales. Extend it (see below) with the names your users actually have. - The gazetteer is fail-open. A Bloom false positive only ever accepts a string (spares a gibberish value) — never wrongly rejects a real one.
- Not a security control. Don't use it to block abuse; it's a UX aid for typo/garbage detection.
Regenerating the artifacts
Both artifacts are committed and reproducible from scratch with a fixed seed.
Corpora
The source name lists are third-party data — git-ignored, not redistributed.
See resources/corpora/SOURCES.md for the exact
sources and licenses (MIT and The Unlicense). Fetch them into resources/corpora/:
Model (resources/data/model.php)
Negatives are synthesized (uniform-random, keyboard walks, letter-shuffled real names),
classes balanced, features standardized, logistic regression fit with L2, and the
threshold calibrated by Youden's J (or --target-fpr=). The calibration test asserts the
held-out FPR ≤ 8% and recall ≥ 90%.
Gazetteer (resources/data/gazetteer.php)
Add multi-locale lists so non-Anglo names are whitelisted — this is the lever that fixes
false positives. A curated supplement ships in resources/names/multilocale.txt.
Regenerating
model.php/gazetteer.phpis a deliberate, reviewed commit — note the corpus and command in the commit body.
Testing
License
The MIT License (MIT). See LICENSE.
All versions of gibberish-name-detector with dependencies
illuminate/contracts Version ^12.0 || ^13.0
illuminate/support Version ^12.0 || ^13.0