Scrapy

From Wikipedia, the free encyclopedia

Template:Short description Script error: No such module "Distinguish".

Page Module:Infobox/styles.css has no content.

Scrapy
[[Programmer|DeveloperTemplate:Pluralize from text]]Zyte (formerly Scrapinghub)
Initial releaseScript error: No such module "Date time".
Template:Infobox software/simple
Written inPython
Operating systemWindows, macOS, Linux
TypeWeb crawler
LicenseBSD License
WebsiteScript error: No such module "WikidataIB".

Script error: No such module "Check for conflicting parameters".

Scrapy (/ˈskrp/[1] Script error: No such module "Respell".) is a free and open-source web-crawling framework written in Python. Originally designed for web scraping, it can also be used to extract data using APIs or as a general-purpose web crawler.[2] It is currently maintained by Zyte (formerly Scrapinghub), a web-scraping development and services company.

Scrapy project architecture is built around "spiders", which are self-contained crawlers that are given a set of instructions. Following the spirit of other don't repeat yourself frameworks, such as Django,[3] it makes it easier to build and scale large crawling projects by allowing developers to reuse their code.

Some well-known companies and products using Scrapy are: Lyst,[4][5] Parse.ly,[6] Sayone Technologies,[7] Sciences Po Medialab,[8] Data.gov.uk’s World Government Data site.[9]

History

Scrapy was born at London-based web-aggregation and e-commerce company Mydeco, where it was developed and maintained by employees of Mydeco and Insophia (a web-consulting company based in Montevideo, Uruguay). The first public release was in August 2008 under the BSD license, with a milestone 1.0 release happening in June 2015.[10] In 2011, Zyte (formerly Scrapinghub) became the new official maintainer.[11][12]

References

Page Template:Reflist/styles.css has no content.

  1. ^ Page Module:Citation/CS1/styles.css has no content."Commit 975f150". GitHub. Archived from the original on 2021-10-18. Retrieved 2021-10-18.
  2. ^ Scrapy at a glance Script error: No such module "webarchive"..
  3. ^ Page Module:Citation/CS1/styles.css has no content."Frequently Asked Questions". Frequently Asked Questions, Scrapy 2.8.0 documentation. Archived from the original on 11 November 2020. Retrieved 28 July 2015.
  4. ^ Page Module:Citation/CS1/styles.css has no content.Bell, Eddie; Heusser, Jonathan. "Scalable Scraping Using Machine Learning". Archived from the original on 4 June 2016. Retrieved 28 July 2015.
  5. ^ Page Module:Citation/CS1/styles.css has no content."Scrapy | Companies using Scrapy". Archived from the original on 2020-11-12. Retrieved 2015-07-28.
  6. ^ Page Module:Citation/CS1/styles.css has no content.Montalenti, Andrew (October 27, 2012). "Web Crawling & Metadata Extraction in Python". Web Crawling & Metadata Extraction in Python - Speaker Deck. Archived from the original on September 19, 2020. Retrieved May 11, 2015.
  7. ^ Page Module:Citation/CS1/styles.css has no content."Scrapy Companies". Scrapy | Companies using Scrapy. Archived from the original on 2020-11-12. Retrieved 2017-11-09.
  8. ^ Page Module:Citation/CS1/styles.css has no content."Hyphe v0.0.0: the first release of our new webcrawler is out!". 17 November 2013. Archived from the original on 2016-06-13. Retrieved 2015-07-28.
  9. ^ Script error: No such module "cite tweet".
  10. ^ Page Module:Citation/CS1/styles.css has no content.Medina, Julia (19 June 2015). "Scrapy 1.0 official release out!". scrapy-users (Mailing list). Archived from the original on 25 January 2010. Retrieved 28 July 2015.
  11. ^ Page Module:Citation/CS1/styles.css has no content.Hoffman, Pablo (2013). List of the primary authors & contributors. Archived from the original on 29 May 2017. Retrieved 18 November 2013.
  12. ^ Interview Scraping Hub Script error: No such module "webarchive"..