Resource Library

Cloudera offers a variety of materials on big data consolidation, storage and processing. The library includes high-level overviews as well as detailed information on Apache Hadoop and the surrounding ecosystem.

  1. /content/cloudera/en/resources/library/recordedwebinar/best-practices-for-the-hadoop-data-warehouse-slides/jcr:content/mainContent/resourcecomponent.img.png/1407188576036.png
    Best Practices for the Hadoop Data Warehouse: EDW 101 for Hadoop Professionals
    • Thursday, May 29 2014
    • Category: Video, Why Consolidation Data Platform, Data processing ETL offload, Presentation Slides
    Dr. Ralph Kimball and Eli Collins describe standard data warehouse best practices in Hadoop and how to implement them within a Hadoop environment. This includes identification of dimensions and facts, managing primary keys, and handling slowly changing dimensions (SCDs) and conformed dimensions.
  2. /content/cloudera/en/resources/library/recordedwebinar/large-scale-machine-learning-with-apache-spark/jcr:content/mainContent/resourcecomponent.img.png/1405383605390.png
    Large Scale Machine Learning with Apache Spark
    • Wednesday, May 21 2014
    • Category: Recorded Webinars, Video, CDH, Predictive modeling, Cyber security, Fraud detection
    Spark offers a number of advantages over its predecessor MapReduce that make it ideal for large-scale machine learning. For example, Spark includes MLLib, a library of machine learning algorithms for large data. The presentation will cover the state of MLLib and the details of some of the scalable algorithms it includes, mainly K-means.
  3. /content/cloudera/en/resources/library/recordedwebinar/sas-and-cloudera--analytics-at-scale/jcr:content/mainContent/resourcecomponent.img.png/1405383569146.png
    SAS® and Cloudera Analytics at Scale and Speed
    • Wednesday, May 07 2014
    • Category: Predictive modeling, Data hub, Business process optimization, Software Vendor (ISV), Video, CDH, Recorded Webinars
    Learn about SAS and Cloudera technical integration, how SAS builds on the enterprise data hub, and SAS In-Memory solutions for Hadoop and machine learning capabilities.
  4. /content/cloudera/en/resources/library/productdemo/sas-and-cloudera-demo/jcr:content/mainContent/resourcecomponent.img.png/1405556337477.png
    SAS and Cloudera Demo
    • Wednesday, May 07 2014
    • Category: Predictive modeling, Software Vendor (ISV), Video, Product Demos
    Watch this demo of SAS Visual Analytics where we explore example data from a potential super market wishing to create a new line of organic products.
  5. HBase Design Patterns @ Yahoo!
    • Monday, May 05 2014
    • Category: Presentation, Video, HBaseCon
    This talk reviews some recurring HBase design patterns at Yahoo! and shares some learnings and experiences.
  6. From MongoDB to HBase in Six Easy Months - Operations Session 6
    • Monday, May 05 2014
    • Category: Presentation, Video, HBaseCon
    Pushing well past MongoDB's limits (2TB data every week) is an interesting exercise in operational frustration. It also severely hampers flexibility of design for new use cases. This talk covers the architectural journey from MongoDB/Redis to HBase at Optimizely -- including the performance, design flexibility, speed of implementation, and other gains made. It also covers the operational setup needed to monitor and maintain the system as well as lessons learned from the migration process itself.
  7. /content/cloudera/en/resources/library/hbasecon2014/the-state-of-hbase-replication/jcr:content/mainContent/resourcecomponent.img.png/1405465805610.png
    The State of HBase Replication - Operations Session 2
    • Monday, May 05 2014
    • Category: Video, HBaseCon, Recorded Webinars
    HBase Replication has come a long way since its inception in HBase 0.89 almost four years ago. Today, master-master and cyclic replication setups are supported; many bug fixes and new features like log compression, per-family peers configuration, and throttling have been added; and a major refactoring has been done. This presentation will recap the work done during the past four years, present a few use cases that are currently in production, and take a look at the roadmap.
  8. Blackbird: Storing Billions of Rows a Couple of Milliseconds Away
    • Monday, May 05 2014
    • Category: HBaseCon, Presentation, Video
    Would you use HBase to make billions of rows available for real-time lookup under 10 ms with 99% guarantee? We, at Rocket Fuel, do just that.
  9. HBase: Just the Basics
    • Monday, May 05 2014
    • Category: Video, Presentation, HBaseCon
    A brief Cliff's Notes-level talk covering architecture, API, and schema design
  10. A Survey of HBase Application Archetypes
    • Monday, May 05 2014
    • Category: HBaseCon, Video, Presentation
    This talk presents these archetypes and others based on a use-case survey of clusters conducted by Cloudera's development, product, and services teams.