Download the PHP package mosaiqo/proofread without Composer

On this page you can find all versions of the php package mosaiqo/proofread. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.

FAQ

After the download, you have to make one include require_once('vendor/autoload.php');. After that you have to import the classes with use statements.

Example:
If you use only one package a project is not needed. But if you use more then one package, without a project it is not possible to import the classes with use statements.

In general, it is recommended to use always a project to download your libraries. In an application normally there is more than one library needed.
Some PHP packages are not free to download and because of that hosted in private repositories. In this case some credentials are needed to access such packages. Please use the auth.json textarea to insert credentials, if a package is coming from a private repository. You can look here for more information.

  • Some hosting areas are not accessible by a terminal or SSH. Then it is not possible to use Composer.
  • To use Composer is sometimes complicated. Especially for beginners.
  • Composer needs much resources. Sometimes they are not available on a simple webspace.
  • If you are using private repositories you don't need to share your credentials. You can set up everything on our site and then you provide a simple download link to your team member.
  • Simplify your Composer build process. Use our own command line tool to download the vendor folder as binary. This makes your build process faster and you don't need to expose your credentials for private repositories.
Please rate this library. Is it a good library?

Informations about the package proofread

Proofread

The only eval package native to the official Laravel AI stack.

Latest Release Tests PHP Laravel

Status: Early development — pre-1.0, API unstable.

What it does

Modern Laravel apps increasingly ship AI agents, prompts, and MCP tools straight to production. The official laravel/ai SDK makes building them easy, but the feedback loop for evaluating them has lived outside the framework in language-agnostic tools like Promptfoo. Proofread brings that loop home: a Laravel-native way to measure whether your agents actually do what you think they do, from Pest, from CI, and from production traffic.

Three things make Proofread different:

Installation

Requires PHP 8.3+ and Laravel 13.x. The dashboard additionally needs livewire/livewire (v3 or v4); without it the rest of the package works and the dashboard routes are simply not registered.

Optional MCP integration (expose eval tools to MCP-compatible editors and assistants):

Publish the config and migrations:

Quick start

1. Define an agent with laravel/ai

2. Write a Pest eval

3. Or wrap it in an EvalSuite and run from Artisan

Core concepts

EvalSuite

A suite bundles a dataset, a subject, and a list of assertions. Extend Mosaiqo\Proofread\Suite\EvalSuite and implement three methods:

See src/Suite/EvalSuite.php for the full contract.

Suites can optionally override setUp() and tearDown() for database-dependent setup, and assertionsFor(array $case) to vary assertions per case based on metadata. Both toPassSuite() in Pest and the evals:run Artisan command drive the full lifecycle automatically.

Subjects

Three shapes are accepted:

Mosaiqo\Proofread\Runner\SubjectResolver normalizes all three into a uniform closure and captures token usage, model, provider, latency, and derived cost into assertion metadata.

Assertions

Assertion Kind Example
ContainsAssertion deterministic ContainsAssertion::make('positive')
RegexAssertion deterministic RegexAssertion::make('/^\d+$/')
LengthAssertion deterministic LengthAssertion::between(5, 200)
CountAssertion deterministic CountAssertion::between(1, 10)
JsonSchemaAssertion deterministic JsonSchemaAssertion::fromAgent(MyAgent::class)
TokenBudget operational TokenBudget::maxTotal(1000)
CostLimit operational CostLimit::under(0.01)
LatencyLimit operational LatencyLimit::under(3000)
Rubric semantic (LLM-as-judge) Rubric::make('polite and concise')
Similar semantic (embeddings) Similar::to('reference text')->minScore(0.8)
HallucinationAssertion semantic (LLM-as-judge) HallucinationAssertion::against($groundTruth)
LanguageAssertion semantic (LLM-as-judge) LanguageAssertion::matches('en')
Trajectory trajectory Trajectory::callsTool('search')
GoldenSnapshot snapshot GoldenSnapshot::fromContext()
StructuredOutputAssertion structured StructuredOutputAssertion::conformsTo(MyAgent::class)
PiiLeakageAssertion safety PiiLeakageAssertion::make()

All assertions live under Mosaiqo\Proofread\Assertions\. Each one returns an AssertionResult with a passed bool, a human-readable reason, an optional numeric score, and arbitrary metadata.

Pest expectations

Expectation Subject Usage
toPassAssertion any output expect($output)->toPassAssertion(ContainsAssertion::make('x'))
toPassEval callable / Agent expect($agent)->toPassEval($dataset, $assertions)
toPassSuite EvalSuite expect($suite)->toPassSuite()
toPassRubric string output expect($output)->toPassRubric('polite tone')
toMatchSchema JSON output expect($json)->toMatchSchema($schemaArray)
toCostUnder EvalRun expect($run)->toCostUnder(0.05)
toMatchGoldenSnapshot any output expect($output)->toMatchGoldenSnapshot()

Both toPassEval and toPassSuite assign the resulting EvalRun to the expectation's ->value after the assertion passes, so callers can chain post-run inspection:

Expectations are loaded via Mosaiqo\Proofread\Testing\expectations.php; wire it up from your tests/Pest.php:

A bundled PHPStan extension teaches static analysis about these dynamic expectations — no stub files to maintain.

Artisan commands

Command Purpose
evals:run {suites*} Run one or more EvalSuite classes. Flags: --persist, --fail-fast, --filter, --junit, --queue, --commit-sha, --fake-judge, --concurrency, --gate-pass-rate, --gate-cost-max.
evals:benchmark {suite} Run a suite N times and report pass-rate variance, duration percentiles, cost, and per-case flakiness. Flags: --iterations, --concurrency, --fake-judge, --flakiness-threshold, --format.
evals:compare {base} {head} Structured diff between two persisted runs. Flags: --format=table\|json\|markdown, --only-regressions, --max-cases, --output.
evals:dataset:diff {dataset} Compare two versions of a dataset. Accepts --base, --head, --format.
evals:providers {suite} Run a MultiSubjectEvalSuite and render a matrix of cases × subjects. Flags: --persist, --commit-sha, --concurrency, --provider-concurrency, --fake-judge, --format.
evals:export {id} Export a persisted run or comparison as self-contained Markdown or HTML. id accepts a ULID, a commit SHA prefix, or latest (resolves to the most recent run). Use --type=comparison to target comparisons. Flags: --format, --output, --type=run\|comparison.
evals:cluster Cluster failures by embedding similarity
evals:cost-simulate {agent} Project cost against alternative models using shadow capture data. Flags: --days, --model, --format.
evals:coverage {agent} {dataset} Analyze dataset coverage against shadow captures using embeddings. Flags: --days, --threshold, --max-captures, --embedding-model, --format.
proofread:lint {agents*} Static analysis of Agent instructions(). Flags: --format, --severity, --with-judge.
shadow:evaluate Evaluate captured shadow traffic against registered assertions
shadow:alert Check pass-rate alerts against thresholds
dataset:generate Generate synthetic cases from a schema via an LLM
dataset:import {file} Import a CSV or JSON file into a PHP dataset file. Flags: --name, --output, --force.
dataset:export {dataset} Export a persisted dataset version as CSV or JSON. Flags: --format, --output, --dataset-version.
proofread:make-suite {name} Scaffold a new EvalSuite. Flag: --multi.
proofread:make-assertion {name} Scaffold a new Assertion class.
proofread:make-dataset {name} Scaffold a dataset PHP file. Flag: --path.

Each command supports --help for the full flag list.

Running cases in parallel

EvalRunner::runSuite($suite, concurrency: 5) executes up to 5 cases in parallel via Laravel's process-based concurrency driver. The evals:run --concurrency=N flag exposes the same capability from the CLI.

Concurrency is beneficial for I/O-bound subjects (LLM and HTTP calls) where parallel wait dominates. For deterministic in-memory subjects the per-task overhead of child-process spawning and closure serialization makes sequential execution faster — leave concurrency at its default of 1 in that case.

Subjects must be serializable. Agent class-string FQCNs and static closures serialize cleanly. Ad-hoc closures that capture test-local state by reference (use (&$x)) cannot cross the process boundary and should only be used with concurrency: 1.

Note: concurrency > 1 is unsafe for subjects that write to SQLite. SQLite serializes writers, so parallel SQLite writes will hit "database is locked" errors. Use concurrency for LLM and HTTP subjects, or subjects that only read from the database.

Shadow evals

Shadow evals let you measure agent quality on real production traffic without affecting the user-facing response path.

The flow is:

  1. Apply Mosaiqo\Proofread\Shadow\EvalShadowMiddleware to routes that call your agent. It samples a configurable fraction of requests, sanitizes PII, and persists a ShadowCapture row.
  2. A queued job (or shadow:evaluate on a schedule) runs the registered assertions for that agent over new captures and records a ShadowEval.
  3. shadow:alert checks the pass rate over a sliding window and dispatches a notification (mail, Slack webhook, etc.) when it drops below the configured threshold.
  4. From the dashboard you can promote any failing capture straight into a regression dataset.

Register the assertions you want to evaluate per agent in a service provider:

And enable shadow capture in config/proofread.php:

See src/Shadow/ for the full surface.

Multi-provider comparison

Compare the same dataset against multiple subjects — typically different models, providers, or prompt variations — in a single invocation. Extend MultiSubjectEvalSuite and declare subjects(): array<string, mixed>:

Run it:

Output is a matrix of cases × subjects with pass/fail status and per-subject aggregate stats. The persisted comparison is browsable at /evals/comparisons/{id} in the dashboard, exportable via evals:export {id}, and queryable via the EvalComparison model.

Helper methods on EvalComparison: bestByPassRate(), cheapest(), fastest(). No opinionated overall winner — the three axes are surfaced separately.

Dashboard

A Livewire-powered dashboard ships with the package at /evals (configurable). It is registered only when livewire/livewire (v3 or v4) is installed:

Routes:

Access is controlled by the viewEvals Gate. The default definition allows the local environment only; override it in your AppServiceProvider:

Dataset versioning

Every persisted run is linked to an EvalDatasetVersion snapshot capturing the exact cases that were evaluated. When a dataset's checksum changes between runs, Proofread automatically records a new version without losing the old one. Use evals:dataset:diff to see what changed between any two versions.

Pricing table

Proofread ships approximate pricing for common models (Claude 4.x, GPT-4o, o1 series, Gemini 1.5, OpenAI embeddings). Pricing covers input, output, cache reads, cache writes, and reasoning tokens. CostLimit, the dashboard cost view, and shadow capture cost tracking all work out of the box for supported models.

Override or extend the table in config/proofread.php:

Any model absent from the table reports a null cost — CostLimit fails closed on missing data rather than silently passing.

MCP integration

When laravel/mcp is installed, Proofread exposes four tools:

Register them from your own Server subclass:

And declare which suites are discoverable via the tool in config/proofread.php:

GitHub Actions

Proofread ships a ready-to-use workflow template. Publish it to your project and customize the suite FQCN:

The workflow lands at .github/workflows/proofread.yml and runs your suites on every PR and push to main. It uploads JUnit XML as an artifact and renders a per-case report directly in the PR via mikepenz/action-junit-report.

Required repository secrets if your suites use LLM-backed assertions (Rubric, Hallucination, Similar, Language):

For deterministic CI runs without real LLM calls, add --fake-judge=pass to the evals:run command in the workflow.

PR comments

evals:compare --format=markdown renders the diff as a PR-friendly Markdown document. Regressions lead, improvements follow, stable cases collapse into a <details> block for readability. Pair it with an --output path and post the result via peter-evans/create-or-update-comment or similar:

The published workflow template includes commented scaffolding for this step. Activate it by implementing your own baseline strategy (artifact, shared DB, or branch comparison) to resolve the two run IDs.

Laravel Telescope integration

If your project has laravel/telescope installed, Proofread automatically records persisted eval runs as Telescope entries. They appear under Events tagged with proofread_eval alongside your queries, jobs, and requests. Filter by the proofread_eval tag (or by dataset:..., suite:..., commit:...) to inspect recent runs without leaving your debugging workflow.

Registration is conditional — Proofread checks for Telescope at boot time and wires up the listener only when it is available. No configuration required.

Laravel Pulse integration

If your project uses laravel/pulse, Proofread automatically records persisted eval runs as Pulse metrics. Each run emits a proofread_eval counter keyed by dataset::passed|failed, a proofread_eval_duration gauge (avg + max, in ms), and a proofread_eval_cost sum (in micro-dollars) when cost data is available. They surface alongside your queries, jobs, and requests in the Pulse dashboard.

Publish the optional dashboard card:

Then add it to your Pulse dashboard view resources/views/vendor/pulse/dashboard.blade.php:

Registration is conditional on Pulse being installed and bound in the container — no action needed if you don't use Pulse.

OpenTelemetry integration

If your project has open-telemetry/api installed (with a configured TracerProvider), Proofread emits a span tree for every persisted eval run:

Attributes are prefixed with proofread.* (run.id, dataset.name, suite.class, run.passed, case.duration_ms, etc.). Spans propagate to whatever exporter the host app has configured — Jaeger, Grafana Tempo, Honeycomb, an OTLP collector, and so on.

Registration is conditional on the OpenTelemetry API being available. Proofread does not require it as a hard dependency; install open-telemetry/api (or the full SDK) in your project and wire a TracerProvider via OpenTelemetry\API\Globals to activate tracing.

Laravel Boost integration

If your project uses laravel/boost, publish Proofread's AI guidelines so Boost-powered editors can generate idiomatic suites, assertions, and tests:

The guidelines land at .ai/guidelines/proofread.md. They cover suite structure, assertion selection, testing patterns, and CLI workflow. If your Boost setup expects a different path, move the file after publishing.

CLI subjects (subscription-friendly evaluation)

Most LLM providers bill per-token via their APIs. If you already pay a flat-rate subscription (Max, Pro, etc.), you may not want to rack up API charges on top of what you are already paying. Proofread lets you evaluate against headless CLI tools that run under your subscription.

Claude Code

The CLI must be installed and authenticated in the environment where evals run. claude --output-format json is parsed to extract the response text plus token usage, cost, and model metadata. LatencyLimit, TokenBudget, and CostLimit assertions work unchanged — tokens come from the CLI's reported usage, and cost is whatever the CLI reports (0 for subscription mode).

Extending for other CLIs

Subclass CliSubject and implement three methods:

Optional overrides: timeout(), workingDirectory(), env(), usesStdin(), estimateTokens().

Limitations

Configuration

Publish the config file:

Key sections (see config/proofread.php for the annotated full file):

Testing the package itself

All three must pass on every commit to main.

Contributing

Contributions are welcome. The development workflow is:

See CLAUDE.md for the full set of conventions.

License

The MIT License (MIT). See LICENSE for details.

Credits


All versions of proofread with dependencies

PHP Build Version
Package Version
Requires php Version ^8.3
illuminate/contracts Version ^13.0
laravel/ai Version ~0.5
opis/json-schema Version ^2.6
spatie/laravel-package-tools Version ^1.92
Composer command for our command line client (download client) This client runs in each environment. You don't need a specific PHP version etc. The first 20 API calls are free. Standard composer command

The package mosaiqo/proofread contains the following files

Loading the files please wait ...