📝 Apache Hadoop
Description
< What is it? >
Apache Hadoop is an open-source framework for distributed storage and processing of large datasets. Its core modules provide shared utilities, a distributed filesystem, cluster resource management, and a batch-processing engine. See the Hadoop module overview.
- Example: store application logs across several machines, allocate resources for a daily processing job, and aggregate the logs into usage statistics.
Key points
< The core modules >
| Module | Full name or role | Responsibility |
|---|---|---|
| Hadoop Common | Shared libraries and utilities | Supports the other Hadoop modules |
| HDFS | Hadoop Distributed File System | Stores files as blocks across machines |
| YARN | Yet Another Resource Negotiator | Schedules applications and manages cluster resources |
| Hadoop MapReduce | Distributed batch-processing engine | Executes mapper and reducer tasks using YARN |
Hadoop names the overall framework. HDFS, YARN, and MapReduce have distinct responsibilities within it.
< How HDFS stores data >
The NameNode manages filesystem metadata, including the namespace and locations of file blocks. DataNodes store the blocks and serve client reads and writes. Clients obtain block locations from the NameNode and transfer file data with DataNodes.
HDFS supports replication to keep copies of blocks on different machines. The replication factor is configurable. Replication helps tolerate machine failures; it does not protect against every cause of data loss, such as an authorized deletion.
Client → NameNode: locate file blocks
Client ↔ DataNodes: read or write block data
HDFS is designed for large files and high-throughput access. Many tiny files create metadata overhead and can reduce efficiency. See the HDFS architecture guide.
< How processing uses the cluster >
A Hadoop MapReduce application reads input, runs map tasks, shuffles records by key, and runs reduce tasks to produce output. YARN allocates the resources used by the application. A processing engine such as Spark can also run on YARN and read HDFS data; storing a file in HDFS does not require processing it with MapReduce.
Comparison
< Hadoop and Spark >
| Aspect | Hadoop | Spark |
|---|---|---|
| Scope | Framework containing storage, resource management, and processing modules | Distributed data-processing engine |
| Storage | HDFS is a core module | Connects to HDFS, object storage, and other data sources |
| Processing | Hadoop MapReduce supplies a batch engine | Supports multi-stage batch computations, SQL, and Structured Streaming |
| Working together | HDFS stores data and YARN allocates resources | Spark performs the computation using those resources |
The processing-engine comparison is Hadoop MapReduce versus Spark. The broader systems can be complementary. See the Spark cluster overview.
Implementation
< Write and read a file in HDFS >
This example assumes an installed Hadoop client, a running HDFS cluster configured as the default filesystem, and a writable HDFS home directory. It does not install or start a cluster. Choose a fresh example directory when rerunning it.
# Create a small local input file.
printf 'red blue red\n' > hadoop-demo-input.txt
# Relative HDFS paths resolve under the current HDFS user's home directory.
hdfs dfs -mkdir -p hadoop-demo
hdfs dfs -put hadoop-demo-input.txt hadoop-demo/input.txt
hdfs dfs -cat hadoop-demo/input.txt
Expected output from the final command:
red blue red
The local file and HDFS file live in different filesystems. -put copies the local file into HDFS; -cat reads the stored file. The filesystem shell reference documents these commands. To process the text, a word-count job could then produce red → 2 and blue → 1, as explained on the MapReduce page.
Troubleshoot
< Common problems >
| Symptom | What to inspect |
|---|---|
| Client cannot reach HDFS | Filesystem configuration, NameNode availability, and network access |
| Permission denied | HDFS user identity, directory ownership, and permissions |
| Upload reports that the destination exists | Use a fresh destination path; -put does not overwrite by default |
| A YARN application waits for resources | Queue capacity, requested memory/CPU, and available cluster resources |
| Many small files slow the system | File layout and NameNode metadata load; consider consolidating files |
Related ideas
- MapReduce explains the map, shuffle, and reduce processing model.
- Apache Spark can process data stored in HDFS and run on YARN.
- Snowflake provides a managed cloud data platform.
- Distributed Systems & Data Processing provides the broader context.