Libraries tagged by crawlora
godbout/htmlpagedom
30339 Downloads
jQuery-inspired DOM manipulation extension for Symfony's Crawler
crossjoin/browscap
234475 Downloads
The standalone PHP Browscap parser Crossjoin\Browscap detects browser properties as well as device information based on the user agent string of the requesting browsers and search engines, using the data from the Browser Capabilities Project. It's several hundred times faster than the build-in PHP function get_browser(), and faster than other Browscap PHP libraries, with much lower memory consumption. Optionally Crossjoin\Browscap automatically updates the Browscap data, so you're always up-to-date. The newest version is build for PHP 7.x, for PHP >= 5.6 use version 2.x, for PHP >= 5.3 use version 1.x.
zenstruck/dom
11634 Downloads
DOM crawler with advanced selector API and assertions.
xcrawler/xcrawler
881 Downloads
A fast, simple and powerful PHP web crawler (scraper/spider) 快速、简洁且强大的爬虫/采集框架
silverstripe-labs/googleanalytics
10290 Downloads
The Google Analytics module consists of 2 components that can be employed independently: The Google Logger injects the google analytics javascript snippet into your source code and logs relevant events (as of now only crawler visits) The Analyzer adds the Google Analytics UI to your CMS.
rtfirst/llms-txt
2167 Downloads
LLMs.txt Generator - Generates llms.txt files for AI/LLM crawlers with website content in Markdown format, with optional API key protection.
octopoda/octopus
5015 Downloads
PHP Sitemap crawler
marcvanh/laravel-bot-block
3986 Downloads
A custom middleware package for Laravel. Temporarily blocks crawlers scanning for vulnerabilities.
johnfmorton/craft-llm-ready
332 Downloads
Serve Markdown versions of your Craft CMS pages to AI crawlers and LLMs via .md URLs, content negotiation, and user-agent detection.
flancer32/mage2_ext_bot_sess
12300 Downloads
Magento2: prevent session creation for bots & crawlers.
bugbuster/contao-botdetection-bundle
47451 Downloads
Contao bundle helper class to detect search engines, bots, spiders, crawlers ...
codeguy/arachnid
9789 Downloads
A crawler to find all unique internal pages on a given website
yurunsoft/crawler
83 Downloads
宇润爬虫框架(Yurun Crawler) 是一个低代码、高性能、分布式爬虫采集框架,这可能是最一把梭的爬虫框架。
yuan1994/z-crawler
53 Downloads
【正方教务】爬虫,支持成绩查询、考试查询、课表查询、四六级成绩查询、四六级报名、选课查询、修改密码、获取用户菜单等功能,并且解析数据成易读格式,符合 psr 规范,拿来即用
webprofil/crawler
1836 Downloads
Crawls Website and can generate backstop json