Download the PHP package commently/crawler without Composer

On this page you can find all versions of the php package commently/crawler. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.

FAQ

After the download, you have to make one include require_once('vendor/autoload.php');. After that you have to import the classes with use statements.

Example:
If you use only one package a project is not needed. But if you use more then one package, without a project it is not possible to import the classes with use statements.

In general, it is recommended to use always a project to download your libraries. In an application normally there is more than one library needed.
Some PHP packages are not free to download and because of that hosted in private repositories. In this case some credentials are needed to access such packages. Please use the auth.json textarea to insert credentials, if a package is coming from a private repository. You can look here for more information.

  • Some hosting areas are not accessible by a terminal or SSH. Then it is not possible to use Composer.
  • To use Composer is sometimes complicated. Especially for beginners.
  • Composer needs much resources. Sometimes they are not available on a simple webspace.
  • If you are using private repositories you don't need to share your credentials. You can set up everything on our site and then you provide a simple download link to your team member.
  • Simplify your Composer build process. Use our own command line tool to download the vendor folder as binary. This makes your build process faster and you don't need to expose your credentials for private repositories.
Please rate this library. Is it a good library?

Informations about the package crawler

Laravel Polite Crawler

A minimal, polite URL crawler for Laravel. Give it a URL and it brings you the raw document — respecting robots.txt, per-host rate limits and conditional requests. It does not know (or care) what it crawls. Parsing and consuming the content is entirely your job, so the package plugs into any project without dictating a schema.

Why this package

Most "crawler" packages are opinionated about what you crawl: they discover links, follow HTML, render JavaScript, or ship their own database schema. laravel-polite-crawler deliberately does none of that. It is just the polite transport layer:

The collection pipeline stays in your application:

Features

Requirements

Installation

If you are developing against a local checkout (the recommended workflow for this package), use a path repository in your composer.json:

The service provider is auto-discovered. Publish the configuration when you want to tune the defaults:

Usage

Fetch a single URL

CrawlRequest is a plain DTO:

Property Description
url The URL to fetch.
key Stable identifier used to key batch results; defaults to the URL.
etag Sent as If-None-Match.
lastModified Sent as If-Modified-Since.
headers Extra request headers merged on top of the defaults.

CrawlResponse carries the raw body and transport metadata:

Property Description
outcome One of CrawlOutcome::Success/NotModified/Error/Throttled/Blocked.
status HTTP status code (0 for transport errors / gated attempts).
body Raw response body (empty for not_modified).
headers Response headers (header(string $name) for case-insensitive lookup).
effectiveUri Final URL after redirects.
duration Request time in seconds, when available.
etag / lastModified Validators returned by the server.
retryAfterSeconds Parsed Retry-After for 429 responses.
error Reason message for error outcomes.

Fetch many URLs concurrently

fetchMany() crawls a list of requests in waves, applying the same politeness rules, and returns only the attempts that were actually fired:

The permit hook

Pass a permit callable to refuse individual requests right before they are fired — useful for app-level idempotency (e.g. "has this feed already been fetched by another worker?"):

Sink + queued job

The package ships a CrawlUrl queued job and a Sink contract. The job fetches the URL (enforcing politeness) and hands the document to whatever Sink is bound in the container:

Bind it in a service provider:

Then dispatch crawls anywhere:

Note: the job only calls the sink for success / not_modified outcomes. error, throttled and blocked attempts are dropped (log or requeue them yourself if you care).

Configuration

All settings are optional; sensible defaults apply.

How politeness works

Honest headers

The crawler sends a fixed, honest User-Agent, an XML-capable Accept list and Accept-Encoding: gzip. It never impersonates a browser, rotates user agents, or sends Sec-Fetch-* headers.

Testing

The package ships no tests of its own; it is covered by the consuming application's test suite (see the HTTP-fake based tests for rate limiting, robots.txt and conditional-request behaviour).

License

MIT


All versions of crawler with dependencies

PHP Build Version
Package Version
Requires php Version ^8.3
illuminate/cache Version ^12.0|^13.0
illuminate/contracts Version ^12.0|^13.0
illuminate/http Version ^12.0|^13.0
illuminate/support Version ^12.0|^13.0
Composer command for our command line client (download client) This client runs in each environment. You don't need a specific PHP version etc. The first 20 API calls are free. Standard composer command

The package commently/crawler contains the following files

Loading the files please wait ...