Download the PHP package bitandblack/document-crawler without Composer
On this page you can find all versions of the php package bitandblack/document-crawler. It is possible to download/install these versions without Composer. Possible dependencies are resolved automatically.
Download bitandblack/document-crawler
More information about bitandblack/document-crawler
Files in bitandblack/document-crawler
Package document-crawler
Short Description Extract different parts of an HTML or XML document.
License MIT
Informations about the package document-crawler
Bit&Black Document Crawler
Extract different parts of an HTML or XML document.
Installation
This library is made for the use with Composer. Add it to your project by running $ composer require bitandblack/document-crawler.
Usage
Using Crawlers to extract parts of a document
The Bit&Black Document Crawler library provides different crawlers, to extract information of a document. There are currently existing:
- AnchorsCrawler: Crawl and extract all defined anchors in a document, that have been declared with
<a href="...">...</a>. - IconsCrawler: Crawl and extract all defined icons in a document, that have been declared with
<link rel="icon" ... />. - ImagesCrawler: Crawl and extract all defined images in a document, that have been declared with
<img ... />. - LanguageCodeCrawler: Crawl and extract the language code of a document, that has been declared with
<html lang="...">. - LinkTagsCrawler: Crawl and extract all link tags of a document, that have been declared with
<link ... />. - MetaTagsCrawler: Crawl and extract all defined meta tags in a document, that have been declared with
<meta ... />. - TitleCrawler: Crawl and extract the title of a document, that has been declared with
<title>...</title>.
All those crawlers work the same — they need a DomCrawler object, that contains the document:
You can create a custom Crawler by implementing the CrawlerInterface.
Handling resources
In same cases, resources are getting crawled, which you may want to handle in a specific way. To achieve this, each crawler makes use of a so-called Resource Handler. There are currently existing:
-
The FileSystemDownloadHandler: This one loads resources and writes them to the file system. There are different Http Clients available to fetch resources:
- The HttpDiscoveryClient is the default one and makes use of whatever library your project uses to download resources.
- The
react/httplibrary and fetches resources asynchronously. - You can — for sure — create a custom Http Client by implementing the HttpClientInterface.
- The PassiveResourceHandler: This handler does nothing and is the default one.
You can create a custom Resource Handler by implementing the ResourceHandlerInterface.
Crawling everything at once
In case you don't want to set up something, there is the HolisticDocumentCrawler, that does all the work for you:
The HolisticDocumentCrawler can also be initialised using the createFromUrl method:
Help
If you have any questions, feel free to contact us under [email protected].
Further information about Bit&Black can be found under www.bitandblack.com.
All versions of document-crawler with dependencies
bitandblack/composer-helper Version ^2.0
bitandblack/pathinfo Version ^1.0
fig/http-message-util Version ^1.0
php-http/discovery Version ^1.0
places2be/locales Version ^3.3
psr/http-client Version ^1.0
psr/http-client-implementation Version *
psr/http-factory-implementation Version *
symfony/css-selector Version ^7.0 || ^8.0
symfony/dom-crawler Version ^7.0 || ^8.0