Download the PHP package nitotm/efficient-language-detector without Composer

On this page you can find all versions of the php package nitotm/efficient-language-detector. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.

FAQ

After the download, you have to make one include require_once('vendor/autoload.php');. After that you have to import the classes with use statements.

Example:
If you use only one package a project is not needed. But if you use more then one package, without a project it is not possible to import the classes with use statements.

In general, it is recommended to use always a project to download your libraries. In an application normally there is more than one library needed.
Some PHP packages are not free to download and because of that hosted in private repositories. In this case some credentials are needed to access such packages. Please use the auth.json textarea to insert credentials, if a package is coming from a private repository. You can look here for more information.

  • Some hosting areas are not accessible by a terminal or SSH. Then it is not possible to use Composer.
  • To use Composer is sometimes complicated. Especially for beginners.
  • Composer needs much resources. Sometimes they are not available on a simple webspace.
  • If you are using private repositories you don't need to share your credentials. You can set up everything on our site and then you provide a simple download link to your team member.
  • Simplify your Composer build process. Use our own command line tool to download the vendor folder as binary. This makes your build process faster and you don't need to expose your credentials for private repositories.
Please rate this library. Is it a good library?

Informations about the package efficient-language-detector

Efficient Language Detector

[![downloads](https://img.shields.io/packagist/dt/nitotm/efficient-language-detector)](https://packagist.org/packages/nitotm/efficient-language-detector) ![supported PHP versions](https://img.shields.io/badge/PHP-%3E%3D%207.4-blue) [![license](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](https://www.apache.org/licenses/LICENSE-2.0) [![supported languages](https://img.shields.io/badge/supported%20languages-60-brightgreen.svg)](#languages)

What is a language detector?
It is a tool that identifies which language a text is written in. For example, detect("Gracias") returns "es" for Spanish.


Efficient language detector (Nito-ELD or ELD) is a fast and accurate natural language detection software, written 100% in PHP, with a speed comparable to fast C++ compiled detectors, and accuracy rivaling the best detectors to date.

It has no dependencies, easy installation, all it's needed is PHP with the mb extension.

ELD is also available in: Javascript (v2), C Library (v3, includes Python package & executable), and an outdated Python (v1) implementation. ELD PHP is v3.

  1. Installation
  2. How to use
  3. Benchmarks
  4. Databases
  5. Testing
  6. Languages

Installation

ELD execution options

Low memory Modes
For Modes string, bytes and disk, all size databases can run with 128MB PHP default setting
disk mode with extralarge size, can run with almost no RAM as it only uses 0.5MB
string and bytes are a great choice for general use, as they are just ~2x slower than array

Fastest Mode: Array (higher memory usage)
For array Mode it is recommended to use OPcache, specially for the larger databases to reduce load times
We need to set opcache.interned_strings_buffer, opcache.memory_consumption high enough for each database

Check Databases for more info.

How to use?

detect() expects a UTF-8 string and returns an object with a language property, containing an ISO 639-1 code (or other selected scheme), or 'und' for undetermined language.

To select database Size and Mode, or language output scheme. We can import Eld... constants to see avalible options.

Languages subsets

Calling langSubset() once, will set the subset.

Other Functions

There is a CLI wrapper (BETA version)
>./bin/eld --help on Linux.
>php bin/eld --help on Windows.

Benchmarks

I compared ELD with a different variety of detectors, as there are not many in PHP.

URL Version Core Language
https://github.com/nitotm/efficient-language-detector/ 3.1.0 PHP
https://github.com/pemistahl/lingua-py 2.0.2 Rust
https://github.com/facebookresearch/fastText 0.9.2 C++
https://github.com/CLD2Owners/cld2 Aug 21, 2015 C++
https://github.com/patrickschur/language-detection 5.3.0 PHP
https://github.com/wooorm/franc 7.2.0 Javascript

Benchmarks:

Time execution benchmark for ELD size large ( check others sizes at more benchmarks )

timetable accuracy table

Databases

Low memory database modes

Modes 'bytes' and 'string' are very similar, they differ on how they are load, and are just 2x slower than Array
Mode 'string' can be OPcache'd, more expensive compilation, but then instant load, 'bytes' has always a steady ~fast load
Special mention to 'disk' mode, while slower, is the fastest uncached load & detect for the larger databases

Mode Disk Bytes String Bytes String
Database Size option Extralarge Extralarge Extralarge Large Large
File size 39 MB 39 MB 39 MB 20 MB 20 MB
Memory usage 0.4 MB 40 MB 40 MB 22 MB 22 MB
Memory usage Cached 0.4 MB 40 MB 0.4 MB + OP 22 MB 0.4 MB + OP
Memory peak 0.4 MB 40 MB 56 MB 22 MB 32 MB
Memory peak Cached 0.4 MB 40 MB 0.4 MB + OP 22 MB 0.4 MB + OP
OPcache used memory - - 39 MB - 20 MB
OPcache used interned - - 0.4 MB - 0.4 MB
Load & detect() Uncached 0.0012 sec 0.04 sec 0.25 sec 0.02 sec 0.11 sec
Load & detect() Cached 0.0011 sec 0.04 sec 0.0003 sec 0.02 sec 0.0003 sec
Mode Bytes String Bytes String
Database Size option Medium Medium Small Small
File size 6 MB 6 MB 2 MB 2 MB
Memory usage 8 MB 8 MB 2 MB 2 MB
Memory usage Cached 8 MB 0.4 MB + OP 2 MB 0.4 MB + OP
Memory peak 8 MB 12 MB 2 MB 3 MB
Memory peak Cached 8 MB 0.4 MB + OP 2 MB 0.4 MB + OP
OPcache used memory - 0 MB - 0 MB
OPcache used interned - 6 MB - 2 MB
Load & detect() Uncached 0.006 sec 0.04 sec 0.003 sec 0.016 sec
Load & detect() Cached 0.006 sec 0.0003 sec 0.002 sec 0.0003 sec

Fastest mode Array, but memory hungry

Array Mode, Size: Small Medium Large Extralarge
Pros Lowest memory Equilibrated Fastest Most accurate
Cons Least accurate Slowest (but fast) High memory Highest memory
File size 3 MB 9 MB 28 MB 64 MB
Memory usage 46 MB 137 MB 547 MB 1143 MB
Memory usage Cached 0.4 MB + OP 0.4 MB + OP 0.4 MB + OP 0.4 MB + OP
Memory peak 78 MB 287 MB 969 MB 2047 MB
Memory peak Cached 0.4 MB + OP 0.4 MB + OP 0.4 MB + OP 0.4 MB + OP
OPcache used memory 21 MB 70 MB 242 MB 516 MB
OPcache used interned 4 MB 10 MB 45 MB 91 MB
Load & detect() Uncached 0.13 sec 0.5 sec 1.4 sec 3.2 sec
Load & detect() Cached 0.0003 sec 0.0003 sec 0.0003 sec 0.0003 sec
Settings (Recommended)
memory_limit >= 128 >= 340 >= 1060 >= 2200
opcache.interned...* >= 8 (16) >= 16 (32) >= 60 (70) >= 116 (128)
opcache.memory >= 64 (128) >= 128 (230) >= 360 (450) >= 750 (820)

Testing

Default composer install might not include these files. Use --prefer-source to include them.

Languages

am, ar, az, be, bg, bn, ca, cs, da, de, el, en, es, et, eu, fa, fi, fr, gu, he, hi, hr, hu, hy, is, it, ja, ka, kn, ko, ku, lo, lt, lv, ml, mr, ms, nl, no, or, pa, pl, pt, ro, ru, sk, sl, sq, sr, sv, ta, te, th, tl, tr, uk, ur, vi, yo, zh

Amharic, Arabic, Azerbaijani (Latin), Belarusian, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Estonian, Basque, Persian, Finnish, French, Gujarati, Hebrew, Hindi, Croatian, Hungarian, Armenian, Icelandic, Italian, Japanese, Georgian, Kannada, Korean, Kurdish (Arabic), Lao, Lithuanian, Latvian, Malayalam, Marathi, Malay (Latin), Dutch, Norwegian, Oriya, Punjabi, Polish, Portuguese, Romanian, Russian, Slovak, Slovene, Albanian, Serbian (Cyrillic), Swedish, Tamil, Telugu, Thai, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Yoruba, Chinese

am, ar, az-Latn, be, bg, bn, ca, cs, da, de, el, en, es, et, eu, fa, fi, fr, gu, he, hi, hr, hu, hy, is, it, ja, ka, kn, ko, ku-Arab, lo, lt, lv, ml, mr, ms-Latn, nl, no, or, pa, pl, pt, ro, ru, sk, sl, sq, sr-Cyrl, sv, ta, te, th, tl, tr, uk, ur, vi, yo, zh

amh, ara, aze, bel, bul, ben, cat, ces, dan, deu, ell, eng, spa, est, eus, fas, fin, fra, guj, heb, hin, hrv, hun, hye, isl, ita, jpn, kat, kan, kor, kur, lao, lit, lav, mal, mar, msa, nld, nor, ori, pan, pol, por, ron, rus, slk, slv, sqi, srp, swe, tam, tel, tha, tgl, tur, ukr, urd, vie, yor, zho


Donations and suggestions

If you wish to donate for open source improvements, hire me for private modifications, request alternative dataset training, or contact me, please use the following link: https://linktr.ee/nitotm


All versions of efficient-language-detector with dependencies

PHP Build Version
Package Version
Requires php Version ^7.4 || ^8.0
ext-mbstring Version *
Composer command for our command line client (download client) This client runs in each environment. You don't need a specific PHP version etc. The first 20 API calls are free. Standard composer command

The package nitotm/efficient-language-detector contains the following files

Loading the files please wait ...