Crawley

Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

[画像:jmg logo]

project.crawley-cloud.com Source Code Docs Changelog

Suggest Changes

Popularity

2.7

Stable

Activity

0.0

Stable

Stars 189

Watchers 18

Forks 34

Last Commit almost 3 years ago

Description

Crawley is a pythonic Scraping / Crawling Framework intended to make easy the way you extract data from web pages into structured storages such as databases.

## Features - High Speed WebCrawler built on Eventlet. - Supports databases engines like Postgre, Mysql, Oracle, Sqlite. - Command line tools. - Extract data using your favourite tool. XPath or Pyquery (A Jquery-like library for python). - Cookie Handlers. - Very easy to use

Programming language: Python

License: GNU General Public License v3.0 only

Tags: Web Crawling Application Frameworks Internet

Latest version: v0.2.1

Crawley alternatives and similar packages

Based on the "Web Crawling" category.
Alternatively, view Crawley alternatives based on common mentions on social networks and blogs.

Scrapy

9.9 9.4 L4 Crawley VS Scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

scrapy logo
pyspider

9.5 0.0 L3 Crawley VS pyspider

DISCONTINUED. A Powerful Spider(Web Crawler) System in Python.

InfluxDB – Built for High-Performance Time Series Workloads

InfluxDB 3 OSS is now GA. Transform, enrich, and act on time series data directly in the database. Automate critical tasks and eliminate the need to move data externally. Download now.

Promo www.influxdata.com

[画像:InfluxDB Logo]

requests-html

9.1 0.0 Crawley VS requests-html

Pythonic HTML Parsing for HumansTM

psf logo
portia

8.8 0.0 L2 Crawley VS portia

Visual scraping for Scrapy

scrapinghub logo
MechanicalSoup

7.7 5.6 L4 Crawley VS MechanicalSoup

A Python library for automating interaction with websites.

MechanicalSoup logo
RoboBrowser

7.2 0.0 L4 Crawley VS RoboBrowser

A simple, Pythonic library for browsing the web without a standalone web browser.

jmcarp logo
Grab

6.4 9.2 L3 Crawley VS Grab

Web Scraping Framework

lorien logo
PSpider

6.4 0.0 Crawley VS PSpider

简单易用的Python爬虫框架,QQ交流群:597510560

xianhu logo
feedparser

6.3 7.5 L3 Crawley VS feedparser

Parse feeds in Python

kurtmckee logo
cola

6.3 0.0 L3 Crawley VS cola

DISCONTINUED. A high-level distributed crawling framework.

qinxuye logo
Scrapely

6.1 0.0 Crawley VS Scrapely

A pure-python HTML screen-scraping library

scrapy logo
gain

6.0 0.0 Crawley VS gain

Web crawling framework based on asyncio.

elliotgao2 logo
Google Search Results in Python

4.2 0.0 Crawley VS Google Search Results in Python

Google Search Results via SERP API pip Python Package

serpapi logo
Sukhoi

4.2 0.0 Crawley VS Sukhoi

Minimalist and powerful Web Crawler.

untwisted logo
MSpider

4.0 0.0 Crawley VS MSpider

Spider

manning23 logo
reader

3.5 9.4 Crawley VS reader

A Python feed reader library.

lemon24 logo
spidy Web Crawler

3.3 0.0 Crawley VS spidy Web Crawler

The simple, easy to use command line web crawler.

rivermont logo
brownant

2.6 0.0 Crawley VS brownant

Brownant is a web data extracting framework.

douban logo
Demiurge

2.2 0.0 L5 Crawley VS Demiurge

PyQuery-based scraping micro-framework.

matiasb logo
Pomp

1.7 0.0 L5 Crawley VS Pomp

Screen scraping and web crawling framework

estin logo
FastImage

1.1 0.0 L4 Crawley VS FastImage

Python library that finds the size / type of an image given its URI by fetching as little as needed

bmuller logo
Mariner

0.5 0.0 Crawley VS Mariner

This a is mirror of Gitlab repository. Open your issues and pull requests there.

radek-sprta logo

* Code Quality Rankings and insights are calculated and provided by Lumnify.
They vary from L1 to L5 with "L5" being the highest.

Do you think we are missing an alternative of Crawley or a related project?

Add another 'Web Crawling' Package

Stream - Scalable APIs for Chat, Feeds, Moderation, & Video.

featured getstream.io

Popular Comparisons

SaaSHub - Software Alternatives and Reviews

featured www.saashub.com

README

Pythonic Crawling / Scraping Framework Built on Eventlet

Build Status Code Climate Stories in Ready

Features

High Speed WebCrawler built on Eventlet.
Supports relational databases engines like Postgre, Mysql, Oracle, Sqlite.
Supports NoSQL databased like Mongodb and Couchdb. New!
Export your data into Json, XML or CSV formats. New!
Command line tools.
Extract data using your favourite tool. XPath or Pyquery (A Jquery-like library for python).
Cookie Handlers.
Very easy to use (see the example).

Documentation

http://packages.python.org/crawley/

Project WebSite

http://project.crawley-cloud.com/

To install crawley run

~$ python setup.py install

or from pip

~$ pip install crawley

To start a new project run

~$ crawley startproject [project_name]
~$ cd [project_name]

Write your Models

""" models.py """
from crawley.persistance import Entity, UrlEntity, Field, Unicode
class Package(Entity):
 #add your table fields here
 updated = Field(Unicode(255)) 
 package = Field(Unicode(255))
 description = Field(Unicode(255))

Write your Scrapers

""" crawlers.py """
from crawley.crawlers import BaseCrawler
from crawley.scrapers import BaseScraper
from crawley.extractors import XPathExtractor
from models import *
class pypiScraper(BaseScraper):
 #specify the urls that can be scraped by this class
 matching_urls = ["%"]
 def scrape(self, response):
 #getting the current document's url.
 current_url = response.url 
 #getting the html table.
 table = response.html.xpath("/html/body/div[5]/div/div/div[3]/table")[0]
 #for rows 1 to n-1
 for tr in table[1:-1]:
 #obtaining the searched html inside the rows
 td_updated = tr[0]
 td_package = tr[1]
 package_link = td_package[0]
 td_description = tr[2]
 #storing data in Packages table
 Package(updated=td_updated.text, package=package_link.text, description=td_description.text)
class pypiCrawler(BaseCrawler):
 #add your starting urls here
 start_urls = ["http://pypi.python.org/pypi"]
 #add your scraper classes here 
 scrapers = [pypiScraper]
 #specify you maximum crawling depth level 
 max_depth = 0
 #select your favourite HTML parsing tool
 extractor = XPathExtractor

Configure your settings

""" settings.py """
import os 
PATH = os.path.dirname(os.path.abspath(__file__))
#Don't change this if you don't have renamed the project
PROJECT_NAME = "pypi"
PROJECT_ROOT = os.path.join(PATH, PROJECT_NAME)
DATABASE_ENGINE = 'sqlite' 
DATABASE_NAME = 'pypi' 
DATABASE_USER = '' 
DATABASE_PASSWORD = '' 
DATABASE_HOST = '' 
DATABASE_PORT = '' 
SHOW_DEBUG_INFO = True

Finally, just run the crawler

~$ crawley run

Do not miss the trending, packages, news and articles with our weekly report.

Awesome Python is part of the LibHunt network. Terms. Privacy Policy.

(CC)

BY-SA

We recommend Spin The Wheel Of Names for a cryptographically secure random name picker.