Skip to content

Latest commit

 

History

History
192 lines (130 loc) · 5 KB

File metadata and controls

192 lines (130 loc) · 5 KB

Newspaper: Article scraping & curation

Release v0.0.5. :ref:`(Installation) <install>`.

Inspired by requests for its simplicity and powered by lxml for its speed.

"Newspaper is an amazing python library for extracting & curating articles." -- tweeted by Kenneth Reitz, Author of requests

"Newspaper delivers Instapaper style article extraction." -- The ChangeLog

We support 10+ languages and everything is in unicode!

>>> import newspaper
>>> newspaper.languages()

Your available langauges are:
input code      full name

  ar              Arabic
  de              German
  en              English
  es              Spanish
  fr              French
  it              Italian
  ko              Korean
  no              Norwegian
  pt              Portuguese
  sv              Swedish
  zh              Chinese

A Glance:

>>> import newspaper

>>> cnn_paper = newspaper.build('http://cnn.com')

>>> for article in cnn_paper.articles:
>>>     print article.url
u'http://www.cnn.com/2013/11/27/justice/tucson-arizona-captive-girls/'
u'http://www.cnn.com/2013/12/11/us/texas-teen-dwi-wreck/index.html'
...

>>> for category in cnn_paper.category_urls():
>>>     print category

u'http://lifestyle.cnn.com'
u'http://cnn.com/world'
u'http://tech.cnn.com'
...
>>> article = cnn_paper.articles[0]
>>> article.download()

>>> article.html
u'<!DOCTYPE HTML><html itemscope itemtype="http://...'
>>> article.parse()

>>> article.authors
[u'Leigh Ann Caldwell', 'John Honway']

>>> article.text
u'Washington (CNN) -- Not everyone subscribes to a New Year's resolution...'

>>> article.top_image
u'http://someCDN.com/blah/blah/blah/file.png'

>>> article.movies
[u'http://youtube.com/path/to/link.com', ...]
>>> article.nlp()

>>> article.keywords
['New Years', 'resolution', ...]

>>> article.summary
u'The study shows that 93% of people ...'

Newspaper has seamless language extraction and detection. If no language is specified, Newspaper will attempt to auto detect a language.

>>> from newspaper import Article
>>> url = 'http://www.bbc.co.uk/zhongwen/simp/chinese_news/2012/12/121210_hongkong_politics.shtml'

>>> a = Article(url, language='zh') # Chinese

>>> a.download()
>>> a.parse()

>>> print a.text[:150]
香港行政长官梁振英在各方压力下就其大宅的违章建
筑(僭建)问题到立法会接受质询,并向香港民众道歉。
梁振英在星期二(12月10日)的答问大会开始之际在其
演说中道歉,但强调他在违章建筑问题上没有隐瞒的意
图和动机。 一些亲北京阵营议员欢迎梁振英道歉,
且认为应能获得香港民众接受,但这些议员也质问梁振英有

>>> print a.title
港特首梁振英就住宅违建事件道歉

If you are certain that an entire news source is in one language, go ahead and use the same api :)

>>> import newspaper
>>> sina_paper = newspaper.build('http://www.sina.com.cn/', langauge='zh')

>>> for category in sina_paper.category_urls():
>>>     print category
u'http://health.sina.com.cn'
u'http://eladies.sina.com.cn'
u'http://english.sina.com'
...

>>> article = sina_paper.articles[0]
>>> article.download()
>>> article.parse()

>>> print article.text
新浪武汉汽车综合 随着汽车市场的日趋成熟,传统的“集
全家之力抱得爱车归”的全额购车模式已然过时,另一种轻
松的新兴 车模式――金融购车正逐步成为时下消费者购买
爱车最为时尚的消费理 念,他们认为,这种新颖的购车模
式既能在短期内
...

>>> print article.title
两年双免0手续0利率 科鲁兹掀背金融轻松购_武汉车市_武汉
汽车网_新浪汽车_新浪网

Features

  • Works in 10+ languages (English, Chinese, German, Arabic, ...)
  • Multi-threaded article download framework
  • News url identification
  • Text extraction from html
  • Top image extraction from html
  • All image extraction from html
  • Keyword extraction from text
  • Summary extraction from text
  • Author extraction from text
  • Google trending terms extraction

User Guide

.. toctree::
   :maxdepth: 2

   user_guide/install
   user_guide/quickstart
   user_guide/advanced

.. toctree::
   :maxdepth: 1

   user_guide/contributors