Resource Library

Cloudera offers a variety of materials on big data consolidation, storage and processing. The library includes high-level overviews as well as detailed information on Apache Hadoop and the surrounding ecosystem.

  1. /content/cloudera/en/resources/library/recordedwebinar/best-practices-for-the-hadoop-data-warehouse-slides/jcr:content/mainContent/resourcecomponent.img.png/1407188576036.png
    Best Practices for the Hadoop Data Warehouse: EDW 101 for Hadoop Professionals
    • Thursday, May 29 2014
    • Category: Video, Why Consolidation Data Platform, Data processing ETL offload, Presentation Slides
    Dr. Ralph Kimball and Eli Collins describe standard data warehouse best practices in Hadoop and how to implement them within a Hadoop environment. This includes identification of dimensions and facts, managing primary keys, and handling slowly changing dimensions (SCDs) and conformed dimensions.
  2. /content/cloudera/en/resources/library/recordedwebinar/large-scale-machine-learning-with-apache-spark-slides/jcr:content/mainContent/resourcecomponent.img.png/1405383623252.png
    Large Scale Machine Learning with Apache Spark
    • Wednesday, May 21 2014
    • Category: CDH, Predictive modeling, Cyber security, Fraud detection, Presentation Slides, Presentation
    Spark offers a number of advantages over its predecessor MapReduce that make it ideal for large-scale machine learning. For example, Spark includes MLLib, a library of machine learning algorithms for large data. The presentation will cover the state of MLLib and the details of some of the scalable algorithms it includes, mainly K-means.
  3. /content/cloudera/en/resources/library/recordedwebinar/large-scale-machine-learning-with-apache-spark/jcr:content/mainContent/resourcecomponent.img.png/1405383605390.png
    Large Scale Machine Learning with Apache Spark
    • Wednesday, May 21 2014
    • Category: Recorded Webinars, Video, CDH, Predictive modeling, Cyber security, Fraud detection
    Spark offers a number of advantages over its predecessor MapReduce that make it ideal for large-scale machine learning. For example, Spark includes MLLib, a library of machine learning algorithms for large data. The presentation will cover the state of MLLib and the details of some of the scalable algorithms it includes, mainly K-means.
  4. /content/cloudera/en/resources/library/productdemo/sas-and-cloudera-demo/jcr:content/mainContent/resourcecomponent.img.png/1405556337477.png
    SAS and Cloudera Demo
    • Wednesday, May 07 2014
    • Category: Predictive modeling, Software Vendor (ISV), Video, Product Demos
    Watch this demo of SAS Visual Analytics where we explore example data from a potential super market wishing to create a new line of organic products.
  5. /content/cloudera/en/resources/library/recordedwebinar/sas-and-cloudera--analytics-at-scale/jcr:content/mainContent/resourcecomponent.img.png/1405383569146.png
    SAS® and Cloudera Analytics at Scale and Speed
    • Wednesday, May 07 2014
    • Category: Predictive modeling, Data hub, Business process optimization, Software Vendor (ISV), Video, CDH, Recorded Webinars
    Learn about SAS and Cloudera technical integration, how SAS builds on the enterprise data hub, and SAS In-Memory solutions for Hadoop and machine learning capabilities.
  6. /content/cloudera/en/resources/library/recordedwebinar/sas-and-cloudera--analytics-at-scale-slides/jcr:content/mainContent/resourcecomponent.img.png/1405383583279.png
    SAS® and Cloudera Analytics at Scale and Speed
    • Wednesday, May 07 2014
    • Category: Predictive modeling, Data hub, Business process optimization, Software Vendor (ISV), CDH, Presentation Slides, Presentation
    Learn about SAS and Cloudera technical integration, how SAS builds on the enterprise data hub, and SAS In-Memory solutions for Hadoop and machine learning capabilities.
  7. /content/cloudera/en/resources/library/whitepaper/why-open-source-matters/jcr:content/mainContent/resourcecomponent.img.png/1405456217973.png
    Cloudera and Open Source
    • Wednesday, May 07 2014
    • File Type: .PDF
    • Category: Document, White Papers, About Hadoop, Open Source Hadoop
    Open source benefits are powerful and time-tested, but they're just “table stakes” when deploying a strategic open source platform like Apache Hadoop -- you'll need so much more from your platform provider.
  8. /content/cloudera/en/resources/library/hbasecon2014/content-identification-using-hbase/jcr:content/mainContent/resourcecomponent.img.jpg/1405466622562.jpg
    Content Identification using HBase
    • Monday, May 05 2014
    • Category: HBaseCon, Presentation, Video
    This presentation will review the options a developer has for HBase querying and retrieval of hash data.
  9. /content/cloudera/en/resources/library/hbasecon2014/hbase-at-bloomberg--high-availability-needs-for-the-financial-in/jcr:content/mainContent/resourcecomponent.img.jpg/1405466661799.jpg
    HBase at Bloomberg: High Availability Needs for the Financial Industry
    • Monday, May 05 2014
    • Category: HBaseCon, Video, Document
    This talk covers data and analytics use cases at Bloomberg and operational challenges around HA. We'll explore the work currently being done under HBASE-10070, further extensions to it, and how this solution is qualitatively different to how failover is handled by Apache Cassandra.
  10. /content/cloudera/en/resources/library/hbasecon2014/opentsdb-2-0-ppt/jcr:content/mainContent/resourcecomponent.img.png/1405466560955.png
    OpenTSDB 2.0
    • Monday, May 05 2014
    • Category: HBaseCon, Presentation Slides
    The OpenTSDB community continues to grow and with users looking to store massive amounts of time-series data in a scalable manner. In this talk, we will discuss a number of use cases and best practices around naming schemas and HBase configuration. We will also review OpenTSDB 2.0's new features, including the HTTP API, plugins, annotations, millisecond support, and metadata, as well as what's next in the roadmap.