Diffbot

From Wikipedia, the free encyclopedia

Template:Short description

Page Module:Infobox/styles.css has no content.Page Template:Infobox company/styles.css has no content.

Diffbot
Company type
Private company
ISINScript error: No such module "WikidataIB".
IndustryScript error: No such module "WikidataIB".
PredecessorScript error: No such module "WikidataIB".
IncorporatedScript error: No such module "WikidataIB".
FoundedScript error: No such module "WikidataIB".
FounderScript error: No such module "WikidataIB".
DefunctScript error: No such module "WikidataIB".
SuccessorScript error: No such module "WikidataIB".
Headquarters,
U.S.
Area served
Worldwide
Key people
Page Template:Plainlist/styles.css has no content.
  • Mike Tung (CEO)
ServicesWeb APIs, Enterprise Search, Web Scraping, Web Crawling
RevenueScript error: No such module "WikidataIB".
Script error: No such module "WikidataIB".
Script error: No such module "WikidataIB".
Total assetsScript error: No such module "WikidataIB".
Number of employees
Script error: No such module "WikidataIB".
ParentScript error: No such module "WikidataIB".
Websitewww.diffbot.com

Script error: No such module "Check for conflicting parameters".Script error: No such module "Check for deprecated parameters".

Diffbot is a developer of machine learning and computer vision algorithms and public APIs for extracting data from web pages / web scraping to create a knowledge base.

Overview

The company has gained interest from its application of computer vision technology to web pages, wherein it visually parses a web page for important elements and returns them in a structured format.[1]

In 2015 Diffbot announced it was working on its version of an automated "knowledge graph" by crawling the web and using its automatic web page extraction to build a large database of structured web data.[2]

In 2019 Diffbot released their Knowledge Graph which has since grown to include over two billion entities (corporations, people, articles, products, discussions, and more), and ten trillion "facts."

Features

The company's products allow software developers to analyze web home pages and article pages,[3] and extract the "important information" while ignoring elements deemed not core to the primary content.[4]

In August 2012 the company released its Page Classifier API, which automatically categorizes web pages into specific "page types".[5] As part of this, Diffbot analyzed 750,000 web pages shared on the social media service Twitter and revealed that photos, followed by articles and videos, are the predominant web media shared on the social network.[6]

In September 2020 the company released a Natural Language Processing API for automatically building Knowledge Graphs from text.[7] [8] The company raised $2 million in funding in May 2012 from investors including Andy Bechtolsheim and Sky Dayton.[9]

Diffbot's customers include Adobe, AOL, Cisco, DuckDuckGo, eBay, Instapaper, Microsoft, Onswipe and Springpad.[4][5][10]

See also

References

Page Template:Reflist/styles.css has no content.

  1. ^ Page Module:Citation/CS1/styles.css has no content."Diffbot Lets Developers Navigate Code the Way Our Eyes See the World". TheNextWeb. August 25, 2011. Retrieved April 21, 2013.
  2. ^ Page Module:Citation/CS1/styles.css has no content."Startup Unleashes Its Clone of Google's 'Knowledge Graph'". Wired. June 4, 2015. Retrieved June 15, 2015.
  3. ^ Page Module:Citation/CS1/styles.css has no content.Kim, Ryan (August 25, 2011). "Diffbot Helps Apps Read the Web Like Humans". Gigaom. Retrieved March 14, 2013.{{cite web}}: CS1 maint: deprecated archival service (link)
  4. ^ a b Page Module:Citation/CS1/styles.css has no content."Investors Back Diffbot's Visual Learning Robot for Web Content". The Wall Street Journal. May 31, 2012. Retrieved March 14, 2013.
  5. ^ a b Page Module:Citation/CS1/styles.css has no content."DiffBot's new API brilliantly reveals what's hiding behind any link". August 16, 2012. Retrieved March 14, 2013.
  6. ^ Page Module:Citation/CS1/styles.css has no content."Twitter: A Day in the Life". Mashable. August 16, 2012. Retrieved March 14, 2013.
  7. ^ Page Module:Citation/CS1/styles.css has no content."New AI Tool Maps the Families of the Bible, A Song of Ice and Fire". Datanami. 2020-09-17. Retrieved 2022-06-08.
  8. ^ Page Module:Citation/CS1/styles.css has no content.Peter, Alex. "Web Scraping". Retrieved 28 March 2021.
  9. ^ Page Module:Citation/CS1/styles.css has no content."Diffbot raises $2 million to help apps understand the open, unstructured web". TheVerge. May 31, 2012. Retrieved March 14, 2013.
  10. ^ Page Module:Citation/CS1/styles.css has no content."Diffbot Bests Google's Knowledge Graph To Feed The Need For Structured Data". Forbes. June 4, 2015. Retrieved June 15, 2015.