Web Crawling and Data Mining with Apache Nutch
By Zakir Laliwala, Abdulbasit Shaikh
Publisher: Packt Publishing
Final Release Date: December 2013
Pages: 136

In Detail

Apache Nutch helps you to create your own search engine and customize it according to your needs. You can integrate Apache Nutch very easily with your existing application and get the maximum benefit from it. It can be easily integrated with different components like Apache Hadoop, Eclipse, and MySQL.

"Web Crawling and Data Mining with Apache Nutch" shows you all the necessary steps to help you in crawling webpages for your application and using them to make your application searching more efficient. You will create your own search engine and will be able to improve your application page rank in searching.

"Web Crawling and Data Mining with Apache Nutch" starts with the basics of crawling webpages for your application. You will learn to deploy Apache Solr on server containing data crawled by Apache Nutch and perform Sharding with Apache Nutch using Apache Solr.

You will integrate your application with databases such as MySQL, Hbase, and Accumulo, and also with Apache Solr, which is used as a searcher.

With this book, you will gain the necessary skills to create your own search engine. You will also perform link analysis and scoring that are helpful in improving the rank of your application page.

Approach

This book is a user-friendly guide that covers all the necessary steps and examples related to web crawling and data mining using Apache Nutch.

Who this book is for

"Web Crawling and Data Mining with Apache Nutch" is aimed at data analysts, application developers, web mining engineers, and data scientists. It is a good start for those who want to learn how web crawling and data mining is applied in the current business world. It would be an added benefit for those who have some knowledge of web crawling and data mining.

Product Details
Recommended for You
Customer Reviews

REVIEW SNAPSHOT®

by PowerReviews
oreillyWeb Crawling and Data Mining with Apache Nutch
 
1.0

(based on 2 reviews)

Ratings Distribution

  • 5 Stars

     

    (0)

  • 4 Stars

     

    (0)

  • 3 Stars

     

    (0)

  • 2 Stars

     

    (0)

  • 1 Stars

     

    (2)

Reviewed by 2 customers

Sort by

Displaying reviews 1-2

Back to top

 
1.0

Not worth reading

By Emir Arnautovic

from Sarajevo, BIH

About Me Developer

Verified Reviewer

Pros

    Cons

      Best Uses

        Comments about oreilly Web Crawling and Data Mining with Apache Nutch:

        This book is poorly written, badly organised, full of incorrect, incomplete and misleading statements, touching variety of topics and technologies, related but not expected to dominate in a book with this title. It is more a set of learning notes of author's first encounter with each of technologies than experts coverage of complex topic.

        Full review is on our blog http://www.atlantbh.com/book-review-web-crawling-and-data-mining-with-apache-nutch/

        (1 of 1 customers found this review helpful)

         
        1.0

        Mostly plagiarized and the rest is poor

        By tdunning

        from Mountain View, CA

        About Me Developer

        Verified Reviewer

        Pros

          Cons

          • Difficult to understand

          Best Uses

            Comments about oreilly Web Crawling and Data Mining with Apache Nutch:

            This book is largely a rehash of the content available on the Nutch wiki, but without giving credit to the original sources at, for example, http://wiki.apache.org/nutch/NutchTutorial.

            What text is added to the freely available content is muddled and difficult to read due to the density of grammatical errors and changes of voice.

            Definitely not recommended. The original text that this book lifted is much more readable.

            Displaying reviews 1-2

            Back to top

             
            Buy 2 Get 1 Free Free Shipping Guarantee
            Buying Options
            Immediate Access - Go Digital what's this?
            Ebook: $20.99
            Formats:  ePub, Mobi, PDF