Download the PHP package tonsoo/php-crawler without Composer

On this page you can find all versions of the php package tonsoo/php-crawler. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.

FAQ

After the download, you have to make one include require_once('vendor/autoload.php');. After that you have to import the classes with use statements.

Example:
If you use only one package a project is not needed. But if you use more then one package, without a project it is not possible to import the classes with use statements.

In general, it is recommended to use always a project to download your libraries. In an application normally there is more than one library needed.
Some PHP packages are not free to download and because of that hosted in private repositories. In this case some credentials are needed to access such packages. Please use the auth.json textarea to insert credentials, if a package is coming from a private repository. You can look here for more information.

  • Some hosting areas are not accessible by a terminal or SSH. Then it is not possible to use Composer.
  • To use Composer is sometimes complicated. Especially for beginners.
  • Composer needs much resources. Sometimes they are not available on a simple webspace.
  • If you are using private repositories you don't need to share your credentials. You can set up everything on our site and then you provide a simple download link to your team member.
  • Simplify your Composer build process. Use our own command line tool to download the vendor folder as binary. This makes your build process faster and you don't need to expose your credentials for private repositories.
Please rate this library. Is it a good library?

Informations about the package php-crawler

Sitemap Generator Crawler

A small, dependency-light PHP crawler that walks a site and generates XML sitemaps. It follows links, respects meta robots directives, and ships with a sitemap extension that can write a single sitemap or rotate into multiple files with an index.

Requirements

Installation

Quick Start

This will crawl https://example.com, write sitemap.xml (or sitemap-2.xml, sitemap-3.xml, etc.), and produce a sitemap-index.xml once multiple sitemap files are created.

Crawler Configuration

The crawler is configured via a fluent API on Crawler:

What these options do

Sitemap Generation

Single sitemap

Rotating sitemap + index

Notes:

Events

You can subscribe to crawler events to observe or extend behavior:

Custom HTTP Client, Logger, and Analyzer

You can plug in your own implementations:

Defaults:

Interfaces to implement:

Error Handling

If maxPages is set and the crawler reaches the limit, it throws LimitExceededException after finishing the crawl loop:

Crawling Behavior

The crawler only processes pages that return an HTML body with a text/html content type. If a page has no HTML body or a non-HTML content type, it is skipped and the corresponding event is emitted.

The crawler collects links from <a href="..."> elements and normalizes them. It will:

This crawler does not parse robots.txt.

Example Script

See examples/crawler.php for a full working example.


All versions of php-crawler with dependencies

PHP Build Version
Package Version
Requires php Version ^8.4
ext-dom Version *
nesbot/carbon Version ^3.11
league/uri Version ^7.8
ext-curl Version *
ext-xmlwriter Version *
Composer command for our command line client (download client) This client runs in each environment. You don't need a specific PHP version etc. The first 20 API calls are free. Standard composer command

The package tonsoo/php-crawler contains the following files

Loading the files please wait ...