This is your work, valued
pig. Mirror of Apache Pig
18idl_storage_guidelines. This document attempts to capture useful patterns and warn about subtle gotchas when it comes to designing and evolving schemas for long-term serialized data. It is not intended as a guide for how to best represent a particular dataset or process.
13elephant-bird. Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, and HBase code.
5piglatin-mode. PigLatin mode for Emacs.
5Vertica-Hadoop-Connector. Vertica Hadoop Connector
2elephant-twin. Elephant Twin is a framework for creating indexes in Hadoop
2elephant-twin-lzo. Elephant Twin LZO uses Elephant Twin to create LZO block indexes
2flume. Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with tunable reliability mechanisms and many failover and recovery mechanisms. The system is centrally managed and allows for intelligent dynamic management. It uses a simple extensible data model that allows for online analytic applications.
1PigEditor. Eclipse plugin for Apache Pig
1hadoop-lzo. Patched, refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20
1awesome-bigdata. A curated list of awesome big data frameworks, ressources and other awesomeness.
1scribe. Scribe is a server for aggregating log data streamed in real time from a large number of servers. It is designed to be scalable, extensible without client-side modification, and robust to failure of the network or any specific machine.
1bud. Prototype Bud runtime (Bloom Under Development)
1giraph. Mirror of Apache Giraph
1