Download the PHP package jcfrane/pdf-text-extractor without Composer
On this page you can find all versions of the php package jcfrane/pdf-text-extractor. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download jcfrane/pdf-text-extractor
More information about jcfrane/pdf-text-extractor
Files in jcfrane/pdf-text-extractor
Package pdf-text-extractor
Short Description A Laravel PDF text extraction package with multiple strategies (PdfParser, XObject, AWS Textract, Tesseract OCR). Handles Canva-generated PDFs, scanned documents, and other edge cases with automatic fallback.
License MIT
Informations about the package pdf-text-extractor
PDF Text Extractor (Laravel)
Laravel-first PDF text extraction with fallback strategies for:
- standard PDFs
- Canva/XObject-based PDFs
- scanned PDFs (via OCR)
Installation
Optional OCR dependencies:
Laravel Setup
The package uses Laravel auto-discovery.
If you want to customize settings, publish config:
This creates:
config/pdf-text-extractor.php
Quick Start (Laravel)
Dependency Injection
Facade
A facade is already included and auto-aliased as PdfTextExtractor.
Configuration
Publish the config file:
This creates config/pdf-text-extractor.php with the following options:
Minimum Text Length
The minimum number of characters an extraction must produce to be considered successful. If a strategy returns fewer characters than this threshold, the next strategy in the list will be tried. Increase this if short garbage output is being accepted; decrease it if your PDFs legitimately contain very little text.
Strategies
An ordered list of extraction strategies. Each strategy is attempted in sequence until one produces text meeting the min_text_length threshold. You can reorder, add, or remove strategies to suit your needs.
| Strategy | Best for | Requirements |
|---|---|---|
PdfParserStrategy |
Standard text-based PDFs | None (included) |
XObjectStrategy |
Canva / XObject-based PDFs | None (included) |
TextractStrategy |
Scanned PDFs (cloud OCR) | aws/aws-sdk-php, AWS credentials |
TesseractStrategy |
Scanned PDFs (local OCR) | tesseract-ocr, ghostscript binaries |
Example: enable all strategies
AWS Textract
Only required if TextractStrategy is in your strategies list. Requires composer require aws/aws-sdk-php.
| Key | Env Variable | Default | Description |
|---|---|---|---|
region |
PDF_EXTRACTOR_AWS_REGION |
us-east-1 |
AWS region for Textract and S3 |
key |
PDF_EXTRACTOR_AWS_KEY |
— | AWS access key ID |
secret |
PDF_EXTRACTOR_AWS_SECRET |
— | AWS secret access key |
version |
PDF_EXTRACTOR_AWS_VERSION |
latest |
AWS SDK version |
s3_bucket |
PDF_EXTRACTOR_AWS_S3_BUCKET |
— | S3 bucket for multi-page PDF processing |
s3_prefix |
PDF_EXTRACTOR_AWS_S3_PREFIX |
pdf-text-extractor |
Key prefix for uploaded PDFs in S3 |
async_poll_interval_ms |
PDF_EXTRACTOR_AWS_ASYNC_POLL_INTERVAL_MS |
1000 |
Milliseconds between polling attempts for async jobs |
async_max_attempts |
PDF_EXTRACTOR_AWS_ASYNC_MAX_ATTEMPTS |
20 |
Maximum number of polling attempts before giving up |
async_delete_uploaded |
PDF_EXTRACTOR_AWS_ASYNC_DELETE_UPLOADED |
true |
Delete the uploaded PDF from S3 after processing |
How Textract works:
- Single-page PDFs use the synchronous
DetectDocumentTextAPI — no S3 required. - Multi-page PDFs use the async flow: the PDF is uploaded to S3,
StartDocumentTextDetectionis called, and the result is polled viaGetDocumentTextDetection.
Add these env values to your .env:
Required IAM permissions:
Tesseract OCR
Only required if TesseractStrategy is in your strategies list. Requires tesseract and ghostscript installed on the system.
| Key | Env Variable | Default | Description |
|---|---|---|---|
binary |
PDF_EXTRACTOR_TESSERACT_BINARY |
tesseract |
Path to the Tesseract binary |
ghostscript_binary |
PDF_EXTRACTOR_GHOSTSCRIPT_BINARY |
gs |
Path to the Ghostscript binary |
language |
PDF_EXTRACTOR_TESSERACT_LANGUAGE |
eng |
Tesseract language code (e.g. eng, fra, deu) |
dpi |
PDF_EXTRACTOR_TESSERACT_DPI |
300 |
DPI used when converting PDF pages to images |
Environment Variables Reference
All env variables at a glance:
Result Object
extract() and extractFromString() return an ExtractionResult:
getText()isSuccessful()getStrategy()getTextLength()
License
MIT
All versions of pdf-text-extractor with dependencies
illuminate/support Version ^11.0|^12.0|^13.0
smalot/pdfparser Version ^2.0