Like the Blog?

Followers

Thursday, 14 September 2017

Setting and Altering Replication & Blocksize in HDFS


1. Configuring Blocksize in HDFS

1.a. Setting the default blocksize for the entire file system

We can set up the default blocksize on HDFS by configuring the hdfs-site.xml configuration file. To do this, we need to add the below property for dfs.blocksize inside the <configuration> tag in hdfs-site.xml (to set the default size as 64MB which is equivalent to 67108864 bytes):

<property>
        <name>dfs.blocksize</name>
        <value>67108864</value>
<description>Default blocksize in HDFS</description>
</property>

We can change this <value> as and when needed and according to our requirement.

To check the default blocksize of the Hadoop File System, we need to run the command below:

hduser@Soumitra-PC:~$ hdfs getconf -confKey dfs.blocksize
67108864


Now if we need to change this default blocksize (to 128MB, equivalent to 134217728 bytes), we can do it by updating the required value dfs.blocksize property in hdfs-site.xml as shown below:

<property>
        <name>dfs.blocksize</name>
        <value>134217728</value>
<description>Default blocksize in HDFS</description>
</property>

We can check the change by the following command again: 

hduser@Soumitra-PC:~$ hdfs getconf -confKey dfs.blocksize
134217728



1.b. Setting the blocksize of a particular file.

If we need to check the blocksize of a particular file in HDFS, the command to do that is as below:

hduser@Soumitra-PC:~$ hdfs dfs -stat %o /file1
67108864

Now, if we want to change this blocksize value to 128MB, i.e, to 134217728 bytes. So let us delete the 'file1' from HDFS directory and re-copy it from the local file system. with the additional 'dfs.blocksize=134217728' pre-mentioned in the command statement itself. The commands are as follows:

hduser@Soumitra-PC:~$ hdfs dfs -rm /file1
17/09/14 22:09:57 INFO fs.TrashPolicyDefault: Namenode trash configuration: Deletion interval = 0 minutes, Emptier interval = 0 minutes.
Deleted /file1

hduser@Soumitra-PC:~$ hdfs dfs -D dfs.blocksize=134217728 -put /home/hduser/file1 /

hduser@Soumitra-PC:~$ hdfs dfs -stat %o /file1
134217728


So, by this way, we can change the blocksize for files on individual basis, instead of changing the default blocksize of the whole file system.


2. Configuring Replication in HDFS

2.a. Setting the default replication for the entire file system

We can set up the default replication on HDFS by configuring the hdfs-site.xml configuration file. To do this, we need to add/modify the below property for dfs.replication inside the <configuration> tag in hdfs-site.xml (to set the default size as 1):

 <property>
  <name>dfs.replication</name>
  <value>1</value>
  <description>Default block replication.
  The actual number of replications can be specified when the file is created.
  The default is used if replication is not specified in create time.
  </description>
 </property>

We can change this <value> as and when needed and according to our requirement.

To check the replication of file(s) in Hadoop File System, we can do a simple ls on the HDFS and check the second column of the outpit, whichs shows the replication of all files in the HDFS:

hduser@Soumitra-PC:~$ hdfs dfs -ls /
Found 6 items
drwxr-xr-x   - hduser supergroup          0 2017-09-12 07:26 /WordCount
-rw-r--r--   1 hduser supergroup         43 2017-09-14 22:11 /file1
-rw-r--r--   1 hduser supergroup         30 2017-09-13 13:42 /file2
drwxr-xr-x   - hduser supergroup          0 2017-09-10 19:41 /hadoop
drwxr-xr-x   - hduser supergroup          0 2017-09-09 21:56 /system
-rw-r--r--   1 hduser supergroup        209 2017-09-14 13:50 /test.cc

Mostly, we can see all files are having replication 1. This may happen because of the default replication set as 1 in hdfs-site.xml file.


Now if we need to change this default replication (to 3), we can do it by updating the required value dfs.replication property in hdfs-site.xml as shown below:

 <property>
  <name>dfs.replication</name>
  <value>3</value>
  <description>Default block replication.
  The actual number of replications can be specified when the file is created.
  The default is used if replication is not specified in create time.
  </description>
 </property>

Now, let's copy a file from local file system to the HDFS home directory and do a ls and check the replication:

hduser@Soumitra-PC:~$ hdfs dfs -put /home/soumitra/test.txt /

hduser@Soumitra-PC:~$ hdfs dfs -ls /
Found 7 items
drwxr-xr-x   - hduser supergroup          0 2017-09-12 07:26 /WordCount
-rw-r--r--   1 hduser supergroup         43 2017-09-14 22:11 /file1
-rw-r--r--   1 hduser supergroup         30 2017-09-13 13:42 /file2
drwxr-xr-x   - hduser supergroup          0 2017-09-10 19:41 /hadoop
drwxr-xr-x   - hduser supergroup          0 2017-09-09 21:56 /system
-rw-r--r--   1 hduser supergroup        209 2017-09-14 13:50 /test.cc
-rw-r--r--   3 hduser supergroup         21 2017-09-14 22:45 /test.txt



2.b. Setting the replication of a particular file.

If we need to change the replication of a particular file in HDFS instead of the default replication of the whole file system. The command to do that is as below:

hduser@Soumitra-PC:~$ hdfs dfs -setrep 2 /test.txt
Replication 2 set: /test.txt

hduser@Soumitra-PC:~$ hdfs dfs -ls /
Found 7 items
drwxr-xr-x   - hduser supergroup          0 2017-09-12 07:26 /WordCount
-rw-r--r--   1 hduser supergroup         43 2017-09-14 22:11 /file1
-rw-r--r--   1 hduser supergroup         30 2017-09-13 13:42 /file2
drwxr-xr-x   - hduser supergroup          0 2017-09-10 19:41 /hadoop
drwxr-xr-x   - hduser supergroup          0 2017-09-09 21:56 /system
-rw-r--r--   1 hduser supergroup        209 2017-09-14 13:50 /test.cc
-rw-r--r--   2 hduser supergroup         21 2017-09-14 22:45 /test.txt

Here, we have changed the replication of test.txt file from 3 to 2 using the setrep command, and displayed the result using ls. The change is clearly reflected in the output.







Document prepared by Mr. Soumitra Ghosh

Assistant Professor, Information Technology,
C.V.Raman College of Engineering, Bhubaneswar
Contact: soumitraghosh@cvrce.edu.in

Various ways of Starting and Stopping Nodes in Hadoop


After you have logged in as the dedicated user for Hadoop(in my case it is hduser) that you must have created while installation, go to the installation folder of Hadoop(in my case it is /usr/local/hadoop). Inside the directory hadoop, there will be a folder 'sbin', where there will be a several files like start-all.sh, stop-all.sh, start-dfs.sh, stop-dfs.sh, hadoop-daemons.sh, yarn-daemons.sh, etc. Executing these files can help us start and/or stop in various ways.

soumitra@Soumitra-PC:~$ su hduser
Password:  

hduser@Soumitra-PC:~$ cd /usr/local/hadoop/sbin

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ ls
distribute-exclude.sh    start-all.cmd        stop-balancer.sh
hadoop-daemon.sh         start-all.sh         stop-dfs.cmd
hadoop-daemons.sh        start-balancer.sh    stop-dfs.sh
hdfs-config.cmd          start-dfs.cmd        stop-secure-dns.sh
hdfs-config.sh           start-dfs.sh         stop-yarn.cmd
httpfs.sh                start-secure-dns.sh  stop-yarn.sh
kms.sh                   start-yarn.cmd       yarn-daemon.sh
mr-jobhistory-daemon.sh  start-yarn.sh        yarn-daemons.sh
refresh-namenodes.sh     stop-all.cmd
slaves.sh                stop-all.sh

1. Starting and Stopping all the components at the same time:

#You can start and stop all the daemons at the same time, by using start-all.sh and stop-all.sh #commands:

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ start-all.sh
This script is Deprecated. Instead use start-dfs.sh and start-yarn.sh
Starting namenodes on [localhost]
localhost: starting namenode, logging to /usr/local/hadoop/logs/hadoop-hduser-namenode-Soumitra-PC.out
localhost: starting datanode, logging to /usr/local/hadoop/logs/hadoop-hduser-datanode-Soumitra-PC.out
Starting secondary namenodes [0.0.0.0]
0.0.0.0: starting secondarynamenode, logging to /usr/local/hadoop/logs/hadoop-hduser-secondarynamenode-Soumitra-PC.out
starting yarn daemons
starting resourcemanager, logging to /usr/local/hadoop/logs/yarn-hduser-resourcemanager-Soumitra-PC.out
localhost: starting nodemanager, logging to /usr/local/hadoop/logs/yarn-hduser-nodemanager-Soumitra-PC.out

#We do a jps to check whether all components are running or  not.
hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
10072 DataNode
11160 Jps
10441 ResourceManager
10281 SecondaryNameNode
9950 NameNode
10559 NodeManager

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ stop-all.sh
This script is Deprecated. Instead use stop-dfs.sh and stop-yarn.sh
Stopping namenodes on [localhost]
localhost: stopping namenode
localhost: stopping datanode
Stopping secondary namenodes [0.0.0.0]
0.0.0.0: stopping secondarynamenode
stopping yarn daemons
stopping resourcemanager
localhost: stopping nodemanager
no proxyserver to stop

#We do a jps to check whether all components have stopped or  not.
hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
11711 Jps





2. Starting a group of nodes among all at the same time :

#Starting Namenode, Datanode and SecondaryNamenode at the same time.


hduser@Soumitra-PC:/usr/local/hadoop/sbin$ start-dfs.sh
Starting namenodes on [localhost]
localhost: starting namenode, logging to /usr/local/hadoop/logs/hadoop-hduser-namenode-Soumitra-PC.out
localhost: starting datanode, logging to /usr/local/hadoop/logs/hadoop-hduser-datanode-Soumitra-PC.out
Starting secondary namenodes [0.0.0.0]
The authenticity of host '0.0.0.0 (0.0.0.0)' can't be established.
ECDSA key fingerprint is SHA256:e9SM2INFNu8NhXKzdX9bOyKIKbMoUSK4dXKonloN7JY.
Are you sure you want to continue connecting (yes/no)? yes
0.0.0.0: Warning: Permanently added '0.0.0.0' (ECDSA) to the list of known hosts.
0.0.0.0: starting secondarynamenode, logging to /usr/local/hadoop/logs/hadoop-hduser-secondarynamenode-Soumitra-PC.out

#Starting ResourceManager daemon and NodeManager daemon:

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ start-yarn.sh
starting yarn daemons
starting resourcemanager, logging to /usr/local/hadoop/logs/yarn-hduser-resourcemanager-Soumitra-PC.out
localhost: starting nodemanager, logging to /usr/local/hadoop/logs/yarn-hduser-nodemanager-Soumitra-PC.out

#We can check if it's really up and running:

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
14306 DataNode
14660 ResourceManager
14505 SecondaryNameNode
14205 NameNode
14765 NodeManager
15166 Jps



3. Starting and Stopping each node individually

#Starting each node separately

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh start namenode
localhost: starting namenode, logging to /usr/local/hadoop/logs/hadoop-hduser-namenode-Soumitra-PC.out

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
12453 Jps

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh start datanode
localhost: starting datanode, logging to /usr/local/hadoop/logs/hadoop-hduser-datanode-Soumitra-PC.out

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
12621 Jps
12543 DataNode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh start secondarynamenode
localhost: starting secondarynamenode, logging to /usr/local/hadoop/logs/hadoop-hduser-secondarynamenode-Soumitra-PC.out

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12752 Jps
12384 NameNode
12709 SecondaryNameNode
12543 DataNode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ yarn-daemons.sh start resourcemanager
localhost: starting resourcemanager, logging to /usr/local/hadoop/logs/yarn-hduser-resourcemanager-Soumitra-PC.out

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
12852 ResourceManager
12709 SecondaryNameNode
13078 Jps
12543 DataNode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ yarn-daemons.sh start nodemanager
localhost: starting nodemanager, logging to /usr/local/hadoop/logs/yarn-hduser-nodemanager-Soumitra-PC.out

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
13298 Jps
12852 ResourceManager
12709 SecondaryNameNode
13179 NodeManager
12543 DataNode


#Stopping each node separately

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ yarn-daemons.sh stop nodemanager
localhost: stopping nodemanager

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
12852 ResourceManager
12709 SecondaryNameNode
13514 Jps
12543 DataNode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ yarn-daemons.sh stop resourcemanager
localhost: stopping resourcemanager

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
12709 SecondaryNameNode
12543 DataNode
13615 Jps

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh stop secondarynamenode
localhost: stopping secondarynamenode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
13705 Jps
12543 DataNode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh stop datanode
localhost: stopping datanode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
12384 NameNode
13792 Jps

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ hadoop-daemons.sh stop namenode
localhost: stopping namenode

hduser@Soumitra-PC:/usr/local/hadoop/sbin$ jps
13885 Jps




Document prepared by Mr. Soumitra Ghosh

Assistant Professor, Information Technology,
C.V.Raman College of Engineering, Bhubaneswar
Contact: soumitraghosh@cvrce.edu.in

Sunday, 10 September 2017

Word Count Program : An example of a basic MapReduce Program


A MapReduce job usually splits the input data-set into independent chunks which are processed by the map tasks in a completely parallel manner. The framework sorts the outputs of the maps, which are then input to the reduce tasks. Typically both the input and the output of the job are stored in a file-system. 
WordCount is a simple application that counts the number of occurrences of each word in a given input set.
The WordCount operation takes place in two stages: 

i) A Mapper phase: Here, the test is tokenized into words and corresponding key value pairs are formed with these words where the key being the word itself and the value being '1'. 

Example: I love Hadoop as much as I love AI 

After the execution of the Map Phase, the output would look like as shown below:
<I,1>
<love,1>
<Hadoop,1>
<as,1>
<much,1>
<as,1>
<I,1>
<love,1>
<AI,1>

ii) A Reducer phase: Here the keys are grouped together and the values for similar keys are summed up.

So after the Reducer phase has completed its execution, the output would look like as shown below:
<I,2>
<love,2>
<Hadoop,1>
<as,2>
<much,1>
<AI,1>

Thus we get the number of occurrence of each word in the input file.

 
1. Write the WordCount.java program and save it your hduser's home directory like this : /home/hduser/WordCount.java

    package MRExample;
  
     import java.io.IOException;
     import java.util.*;
  
     import org.apache.hadoop.fs.Path;
     import org.apache.hadoop.conf.*;
     import org.apache.hadoop.io.*;
     import org.apache.hadoop.mapred.*;
     import org.apache.hadoop.util.*;
  
     public class WordCount
    {
  
        public static class Map extends MapReduceBase implements Mapper<LongWritable, Text, Text, IntWritable>
       { 
//hadoop supported data types
          private final static IntWritable one = new IntWritable(1);
          private Text word = new Text();
 //map method that performs the tokenizer job and framing the initial key value pairs
          public void map(LongWritable key, Text value, OutputCollector<Text, IntWritable> output, Reporter reporter) throws IOException
         { 
//taking one line at a time and tokenizing the same
            String line = value.toString();
            StringTokenizer tokenizer = new StringTokenizer(line);
//iterating through all the words available in that line and forming the key value pair
            while (tokenizer.hasMoreTokens())
           {
              word.set(tokenizer.nextToken());
//sending to output collector which in turn passes the same to reducer
              output.collect(word, one);
            }
          }
        }
  
        public static class Reduce extends MapReduceBase implements Reducer<Text, IntWritable, Text, IntWritable>
       {
//reduce method accepts the Key Value pairs from mappers, do the aggregation based on keys and produce the final out put

          public void reduce(Text key, Iterator<IntWritable> values, OutputCollector<Text, IntWritable> output, Reporter reporter) throws IOException
         {
            int sum = 0;
//iterates through all the values available with a key and add them together and give the final result as the key and sum of its values
            while (values.hasNext())
            {
              sum += values.next().get();
            }
            output.collect(key, new IntWritable(sum));
          }
        }
  
        public static void main(String[] args) throws Exception
       { 
//creating a JobConf object and assigning a job name for identification purposes
          JobConf conf = new JobConf(WordCount.class);
          conf.setJobName("wordcount");
//Setting configuration object with the Data Type of output Key and Value
          conf.setOutputKeyClass(Text.class);
          conf.setOutputValueClass(IntWritable.class);
//Providing the mapper and reducer class names
          conf.setMapperClass(Map.class);
          conf.setCombinerClass(Reduce.class);
          conf.setReducerClass(Reduce.class);
//sets the FileInputFormat and FileOutputFormat types
          conf.setInputFormat(TextInputFormat.class);
          conf.setOutputFormat(TextOutputFormat.class);
//the hdfs input and output directory to be fetched from the command line  
          FileInputFormat.setInputPaths(conf, new Path(args[0]));
          FileOutputFormat.setOutputPath(conf, new Path(args[1]));
  
          JobClient.runJob(conf);
        }
     }


2. Create Sample Input Files (say, file1, file2 in the path /home/hduser/ as before, i.e, in your local file system)

file1 : I watch Game of Thrones. I Watch House MD.
file2 : I love to watch Sherlock too.

hduser@Soumitra-PC:~$ cat file1
I watch Game of Thrones. I Watch House MD.
hduser@Soumitra-PC:~$ cat file2
I love to watch Sherlock too.


3. Make a directory (say, MRClasses in the same path /home/hduser) which will contain the class files(WordCount Class, Map Class and Reduce Class file) after the WordCount.java is compiled: 

hduser@Soumitra-PC:~$ mkdir MRClasses


I have done a ls and checked that the directory MRClasses has been indeed created or not, along with the files file1 and file2 that we have created in step 2.

4. Make a directory(say, WordCount) in HDFS to store the WordCount Program's Input and hold the Program's output. 
For that we create a directory WordCount, and inside that I have created another directory(/WordCount/Input) to store my program's input (i.e, file1 and file2)

hduser@Soumitra-PC:~$ hdfs dfs -mkdir /WordCount /WordCount/Input
 

5. Copy the Input Files from local file system to HDFS (i.e, copy from /home/hduser to /WordCount/Input): 

hduser@Soumitra-PC:~$ hdfs dfs -copyFromLocal /home/hduser/file1 /home/hduser/file2 /WordCount/Input 

We can do a ls to check whether the directories are correctly created or not. 
 

Also, check the contents of the files after copying into HDFS:

hduser@Soumitra-PC:~$ hdfs dfs -copyFromLocal /home/hduser/file /home/hduser/file2 /WordCount/Input

We can do a cat command on the copied files in HDFS to verify whether the file contents are same as we had created in the local file system.  


6. Run the following command to compile the WordCount.java program. This step will create individual classes: WordCount class, Map Class and a Reduce class and store them in the directory MRClasses.  

hduser@Soumitra-PC:~$ sudo javac -classpath /usr/local/hadoop/share/hadoop/common/hadoop-common-2.6.0.jar:/usr/local/hadoop/share/hadoop/common/lib/hadoop-annotations-2.6.0.jar:/usr/local/hadoop/share/hadoop/mapreduce/hadoop-mapreduce-client-core-2.6.0.jar -d /home/hduser/MRClasses /home/hduser/WordCount.java




7. Create a JAR file out of the classes produced in the last step.
hduser@Soumitra-PC:~$ jar -cvf /home/hduser/WordCount.jar -C /home/hduser/MRClasses/ .



8. Execute the JAR file to run the WordCount program.

hduser@Soumitra-PC:~$ hadoop jar WordCount.jar MRExample.WordCount /WordCount/Input /WordCount/Output


hduser@Soumitra-PC:~$ hdfs dfs -ls /WordCount/Output
hduser@Soumitra-PC:~$ hdfs dfs -cat /WordCount/Output/part-00000




Document prepared by Mr. Soumitra Ghosh

Assistant Professor, Information Technology,
C.V.Raman College of Engineering, Bhubaneswar
Contact: soumitraghosh@cvrce.edu.in

Friday, 8 September 2017

Frequently used Hadoop / HDFS Shell Commands


Here we will discuss some frequently used Hadoop Distributed File System Shell commands that will help us perform many operations on managing files on HDFS.

Most of the commands here are similar to corresponding Unix commands, with respect to syntax as well as their functionality. Only, we need to prefix "hadoop fs -" to the Unix Command. E.g.: hadoop fs -<UNIX-Command>".

Error information is sent to stderr and the output is sent to stdout. So, let's get started.

FS relates to a generic file system which can point to any file systems like local, HDFS etc. But dfs is very specific to HDFS. So when we use FS it can perform operation with from/to local or hadoop distributed file system to destination. But specifying DFS operation relates to HDFS.




But, before we start with executing any Hadoop Command, we need to start our Hadoop on our machine, logged in as the dedicated user that one must have created while installing Hadoop in his/her system.

In my case, the dedicated user created is 'hduser'. So, we need to do 'su hduser' and go to the directory where we have installed hadoop (in my case /usr/local/hadoop), and from 'sbin' folder, we need to execute 'start-all.sh' to start all necessary components in Hadoop. Do a jps to be sure whether all the components are running or not. Please refer the screenshot below:




1) version Check : To check the version of Hadoop.

hduser@Soumitra-PC:/sbin$ hadoop version

2) mkdir and ls Command : HDFS Command to create the directory in HDFS.

#Creating multiple directories at the same time
hduser@Soumitra-PC:/sbin$ hdfs dfs -mkdir /hadoop /soumitra
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /

#Creating sub-directory(sub_folder) inside another directory(soumitra)

hduser@Soumitra-PC:/sbin$ hdfs dfs -mkdir /soumitra/sub_folder
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /soumitra

3) put and copyFromLocal Command : 

put : Copy file from single src, or multiple srcs from local file system to the destination file system.

hduser@Soumitra-PC:/sbin$ hdfs dfs -put /home/soumitra/simple.java /
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /
copyFromLocal : Like 'put' command, it also copies file(s) from Local file system to HDFS.
hduser@Soumitra-PC:/sbin$ hdfs dfs -copyFromLocal /home/soumitra/simple.java /hadoop
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /hadoop

4) get and copyToLocal Command : 

get : Copy file from single src, or multiple srcs from hadoop file system to the local file system.

hduser@Soumitra-PC:/sbin$ hdfs dfs -get /file1 /home/hduser
hduser@Soumitra-PC:/sbin$ ls


copyToLocal : Like 'get' command, it also copies file(s) from HDFS to local FS.
hduser@Soumitra-PC:/sbin$ hdfs dfs -copyToLocal /file1 /home/hduser
hduser@Soumitra-PC:/sbin$ ls



5) df Command : Displays free space at given hdfs destination

hduser@Soumitra-PC:/sbin$ hdfs dfs -df hdfs:/
6) count Command : Count the number of directories, files and bytes under the paths that match the specified file pattern.

hduser@Soumitra-PC:/sbin$ hdfs dfs -count hdfs:/

7) fsck Command : HDFS Command to check the health of the Hadoop file system.

hduser@Soumitra-PC:/sbin$ hdfs fsck - /


8) balancer Command : Run a cluster balancing utility.

hduser@Soumitra-PC:/sbin$ hdfs balancer



9) du Command : Displays size of files and directories contained in the given directory or the size of a file if its just a file.

hduser@Soumitra-PC:/sbin$ hdfs dfs -du /


10) rm Command : HDFS Command to remove the file from HDFS.

hduser@Soumitra-PC:/sbin$ hdfs dfs -rm /hadoop/simple.java


11) rm -r Command : HDFS Command to remove the entire directory and all of its content from HDFS.

hduser@Soumitra-PC:/sbin$ hdfs dfs -rm -r /soumitra




12) expunge Command : HDFS Command that makes the trash empty.


hduser@Soumitra-PC:/sbin$ hdfs dfs -expunge


13) touchz Command : HDFS Command to create a file in HDFS with file size 0 bytes.

hduser@Soumitra-PC:/sbin$ hdfs dfs -touchz /empty_file
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /


14) chmod Command : Change the permissions of files.

hduser@Soumitra-PC:/sbin$ hdfs dfs -chmod 777 /empty_file
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /


15) cat Command : HDFS Command that copies source paths to stdout.

hduser@Soumitra-PC:/sbin$ hdfs dfs -cat /simple.java


16) text Command : HDFS Command that takes a source file and outputs the file in text format.

hduser@Soumitra-PC:/sbin$ hdfs dfs -text /simple.java


17) mv Command : HDFS Command to move files from source to destination. This command allows multiple sources as well, in which case the destination needs to be a directory.

hduser@Soumitra-PC:/sbin$ hdfs dfs -mv /simple.java /soumitra
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /soumitra


18) cp Command : HDFS Command to copy files from source to destination. This command allows multiple sources as well, in which case the destination must be a directory.

hduser@Soumitra-PC:/sbin$ hdfs dfs -cp /soumitra/simple.java /
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /


19) tail Command : Displays last kilobyte of the file "new" to stdout.

hduser@Soumitra-PC:/sbin$ hdfs dfs -tail /simple.java


20) chown Command : HDFS command to change the owner of files.

hduser@Soumitra-PC:/sbin$ hdfs dfs -chown ubuntu:hadoop /empty_file
hduser@Soumitra-PC:/sbin$ hdfs dfs -ls /


21) setrep Command : Default replication factor to a file is 3. Below HDFS command is used to change replication factor of a file.

hduser@Soumitra-PC:/sbin$ hdfs dfs -setrep 3 /simple.java


22) stat Command : Print statistics about the file/directory at <path> in the specified format. Format accepts filesize in blocks (%b), type (%F), group name of owner (%g), name (%n), block size (%o), replication (%r), user name of owner(%u), and modification date (%y, %Y). %y shows UTC date as “yyyy-MM-dd HH:mm:ss” and %Y shows milliseconds since January 1, 1970 UTC. If the format is not specified, %y is used by default.

hduser@Soumitra-PC:/sbin$ hdfs dfs -stat "%F %u:%g %b %o %y %n" /hadoop


23) getfacl Command : Displays the Access Control Lists (ACLs) of files and directories. If a directory has a default ACL, then getfacl also displays the default ACL.

hduser@Soumitra-PC:/sbin$ hdfs dfs -getfacl /hadoop


24) du -s Command : Displays a summary of file lengths.

hduser@Soumitra-PC:/sbin$ hdfs dfs -du -s /simple.java


25) checksum Command : Returns the checksum information of a file.

hduser@Soumitra-PC:/sbin$ hdfs dfs -checksum /simple.java



26) moveFromLocal Command : 

moveFromLocal : Moves file(s) from single source, or multiple sources from 
the local destination file system to HDFS.

hduser@Soumitra-PC:/sbin$ hdfs dfs -moveFromLocal /home/hduser/test.cc /
hduser@Soumitra-PC:/sbin$ ls


27) moveToLocal Command : 

moveToLocal : The command is not implemented yet.
hduser@Soumitra-PC:/sbin$ hdfs dfs -moveToLocal /test.cc /home/hduser






Document Created by Mr. Soumitra Ghosh

Assistant Professor, Information Technology,
C.V.Raman College of Engineering, Bhubaneswar

Contact: soumitraghosh@cvrce.edu.in