Skip to main content

📝 Apache Hadoop

Description​

< What is it? >​

Apache Hadoop is an open-source framework for distributed storage and processing of large datasets. Its core modules provide shared utilities, a distributed filesystem, cluster resource management, and a batch-processing engine. See the Hadoop module overview.

  • Example: store application logs across several machines, allocate resources for a daily processing job, and aggregate the logs into usage statistics.

Key points​

< The core modules >​

ModuleFull name or roleResponsibility
Hadoop CommonShared libraries and utilitiesSupports the other Hadoop modules
HDFSHadoop Distributed File SystemStores files as blocks across machines
YARNYet Another Resource NegotiatorSchedules applications and manages cluster resources
Hadoop MapReduceDistributed batch-processing engineExecutes mapper and reducer tasks using YARN

Hadoop names the overall framework. HDFS, YARN, and MapReduce have distinct responsibilities within it.

< How HDFS stores data >​

The NameNode manages filesystem metadata, including the namespace and locations of file blocks. DataNodes store the blocks and serve client reads and writes. Clients obtain block locations from the NameNode and transfer file data with DataNodes.

HDFS supports replication to keep copies of blocks on different machines. The replication factor is configurable. Replication helps tolerate machine failures; it does not protect against every cause of data loss, such as an authorized deletion.

Client → NameNode: locate file blocks
Client ↔ DataNodes: read or write block data

HDFS is designed for large files and high-throughput access. Many tiny files create metadata overhead and can reduce efficiency. See the HDFS architecture guide.

< How processing uses the cluster >​

A Hadoop MapReduce application reads input, runs map tasks, shuffles records by key, and runs reduce tasks to produce output. YARN allocates the resources used by the application. A processing engine such as Spark can also run on YARN and read HDFS data; storing a file in HDFS does not require processing it with MapReduce.

Comparison​

< Hadoop and Spark >​

AspectHadoopSpark
ScopeFramework containing storage, resource management, and processing modulesDistributed data-processing engine
StorageHDFS is a core moduleConnects to HDFS, object storage, and other data sources
ProcessingHadoop MapReduce supplies a batch engineSupports multi-stage batch computations, SQL, and Structured Streaming
Working togetherHDFS stores data and YARN allocates resourcesSpark performs the computation using those resources

The processing-engine comparison is Hadoop MapReduce versus Spark. The broader systems can be complementary. See the Spark cluster overview.

Implementation​

< Write and read a file in HDFS >​

This example assumes an installed Hadoop client, a running HDFS cluster configured as the default filesystem, and a writable HDFS home directory. It does not install or start a cluster. Choose a fresh example directory when rerunning it.

# Create a small local input file.
printf 'red blue red\n' > hadoop-demo-input.txt

# Relative HDFS paths resolve under the current HDFS user's home directory.
hdfs dfs -mkdir -p hadoop-demo
hdfs dfs -put hadoop-demo-input.txt hadoop-demo/input.txt
hdfs dfs -cat hadoop-demo/input.txt

Expected output from the final command:

red blue red

The local file and HDFS file live in different filesystems. -put copies the local file into HDFS; -cat reads the stored file. The filesystem shell reference documents these commands. To process the text, a word-count job could then produce red → 2 and blue → 1, as explained on the MapReduce page.

Troubleshoot​

< Common problems >​

SymptomWhat to inspect
Client cannot reach HDFSFilesystem configuration, NameNode availability, and network access
Permission deniedHDFS user identity, directory ownership, and permissions
Upload reports that the destination existsUse a fresh destination path; -put does not overwrite by default
A YARN application waits for resourcesQueue capacity, requested memory/CPU, and available cluster resources
Many small files slow the systemFile layout and NameNode metadata load; consider consolidating files

Reference​