خرید بک لینک

Vote count: 0

i need to crawl and extract data for some websites (scraping). the guys at scrapinghub.com have developed:

  • Scrapy: web scraping framework (open-source)
  • Frontera: platform to run a scrapy scrapper in a distributed maer using HBase (open-source)
  • Crawlera: a payed proxy curator and rotator service
  • Scrapy Cloud: a payed production environment to deploy and run and monitor crawls based on scrapy, frontera, crawlera and more

It seems very convinient, but they use python. Is there an equivalent for the Java/Scala world?

I intend to use:

  • jsoup for html parsing,
  • akka for queuing a list of urls to download/parse, in a distributed maer
  • Crawlera for proxy
  • amazon ec2 for deploying

I think that yarn, nutch and spark are an overkill for this application.

do you have any other advice?

asked 36 secs ago

برچسب: نویسنده: استخدام کار تاريخ: شنبه 6 شهريور 1395 ساعت: 6:45

صفحه بندی