hadoop分布式系统

一、hadoop

1.hadoop简介

Hadoop是一个由Apache基金会所开发的分布式系统基础架构。
Hadoop的框架最核心的设计就是:

  • HDFS
  • MapReduce

2.hadoop的技术主要模块

1)Hdfs主要模块:

  • NameNode:是整个文件系统的管理节点。维护整个文件系统的文件目录树,文件/目录的元数据和
    每个文件对应的数据块列表。接收用户的请求。
  • DataNode:是HA(高可用性)的一个解决方案,是备用镜像,但不支持热备

HDFS为海量的数据提供了存储,而MapReduce则为海量的数据提供了计算。
2)Yarn主要模块:

  • ResourceManager
  • NodeManager

3.hadoop的优点

  • 高可靠性。Hadoop按位存储和处理数据的能力值得人们信赖。
  • 高扩展性。Hadoop是在可用的计算机集簇间分配数据并完成计算任务的,这些集簇可以方便地扩展到数以千计的节点中。
  • 高效性。Hadoop能够在节点之间动态地移动数据,并保证各个节点的动态平衡,因此处理速度非常快。
  • 高容错性。Hadoop能够自动保存数据的多个副本,并且能够自动将失败的任务重新分配。
  • 低成本。与一体机、商用数据仓库以及QlikView、Yonghong Z-Suite等数据集市相比,hadoop是开源的,项目的软件成本因此会大大降低。

4.hadoop大数据处理的意义

Hadoop得以在大数据处理应用中广泛应用得益于其自身在数据提取、变形和加载(ETL)方面上的天然优势。Hadoop的分布式架构,将大数据处理引擎尽可能的靠近存储,对例如像ETL这样的批处理操作相对合适,因为类似这样操作的批处理结果可以直接走向存储。Hadoop的MapReduce功能实现了将单个任务打碎,并将碎片任务(Map)发送到多个节点上,之后再以单个数据集的形式加载(Reduce)到数据仓库里。

二、部署hadoop

linux主机 ip
hadoop1 172.25.1.1
hadoop2 172.25.1.2
hadoop3 172.25.1.3
hadoop4 172.25.1.4

软件 =====> 点击下载 提取码: 6br1

1.搭建单机版hadoop

1.建立用户,设置密码

[[email protected] ~]# useradd -u 1000 hadoop
[[email protected] ~]# passwd hadoop

2.hadoop的安装配置
此时需要切换到hadoop用户操作,把需要用的压缩包放到hadoop家目录下

[[email protected] ~]# ls
hadoop-3.0.3.tar.gz  jdk-8u181-linux-x64.tar.gz
[[email protected] ~]# mv * /home/hadoop/
[[email protected] ~]# su - hadoop
Last login: Mon Apr 15 08:46:38 EDT 2019 on pts/0
[[email protected] ~]$ ls
hadoop-3.0.3.tar.gz  jdk-8u181-linux-x64.tar.gz
[[email protected] ~]$ tar zxf jdk-8u181-linux-x64.tar.gz
[[email protected] ~]$ ln -s jdk1.8.0_181/ java
[[email protected] ~]$ tar zxf hadoop-3.0.3.tar.gz
[[email protected] ~]$ ln -s hadoop-3.0.3 hadoop
[[email protected] ~]$ ls
hadoop        hadoop-3.0.3.tar.gz  jdk1.8.0_181
hadoop-3.0.3  java                 jdk-8u181-linux-x64.tar.gz

3.配置环境变量

[[email protected] ~]$ cd hadoop/etc/hadoop/
[[email protected] hadoop]$ vim hadoop-env.sh 
 54 export JAVA_HOME=/home/hadoop/java
 [[email protected] ~]$ vim .bash_profile
PATH=$PATH:$HOME/.local/bin:$HOME/bin:$HOME/java/bin
[[email protected] ~]$ source .bash_profile
[[email protected] ~]$ jps
11629 Jps

4.测试

[[email protected] hadoop]$ pwd
/home/hadoop/hadoop
[[email protected] hadoop]$ mkdir input
[[email protected] hadoop]$ cp etc/hadoop/*.xml input/
[[email protected] hadoop]$ bin/hadoop jar share/hadoop/mapreduce/hadoop-mapreduce-examples-3.0.3.jar grep input output 'dfs[a-z.]+'
	File Input Format Counters 
		Bytes Read=123
	File Output Format Counters 
		Bytes Written=23																					##这就表示成功了
[[email protected] hadoop]$ cd output/
[[email protected] output]$ ls
part-r-00000  _SUCCESS
[[email protected] output]$ cat *
1	dfsadmin

2.搭建伪分布式hadoop

1.编辑配置文件

[[email protected] hadoop]$ pwd
/home/hadoop/hadoop/etc/hadoop
[[email protected] hadoop]$ vim core-site.xml 
<configuration>
    <property>
        <name>fs.defaultFS</name>
        <value>hdfs://localhost:9000</value>
    </property>
</configuration>
[[email protected] hadoop]$ vim hdfs-site.xml
<configuration>
    <property>
        <name>dfs.replication</name>
        <value>1</value>     ##自己充当节点
    </property>
</configuration>

2.生成**做ssh免密

[[email protected] hadoop]$ ssh-******
[[email protected] hadoop]$ ssh-copy-id localhost

3.格式化,并开启服务

[[email protected] hadoop]$ bin/hdfs namenode -format
[[email protected] hadoop]$ cd sbin/
[[email protected] sbin]$ ./start-dfs.sh 
Starting namenodes on [localhost]
Starting datanodes
Starting secondary namenodes [hadoop1]
hadoop1: Warning: Permanently added 'hadoop1,fe80::5054:ff:fe9a:f1f5%eth0' (ECDSA) to the list of known hosts.

4.浏览器查看http://172.25.1.1:9870
hadoop分布式系统
5.测试并上传

[[email protected] hadoop]$ bin/hdfs dfs -mkdir -p /user/hadoop
[[email protected] hadoop]$ bin/hdfs dfs -ls
[[email protected] hadoop]$ bin/hdfs dfs -put input
[[email protected] hadoop]$ bin/hdfs dfs -ls
Found 1 items
drwxr-xr-x   - hadoop supergroup          0 2019-04-15 09:20 input

hadoop分布式系统
删除input和output文件,重新执行命令

[[email protected] hadoop]$ rm -fr input/ output/
[[email protected] hadoop]$ bin/hadoop jar share/hadoop/mapreduce/hadoop-mapreduce-examples-3.0.3.jar grep input output 'dfs[a-z.]+'
[[email protected] hadoop]$ ls
bin  include  libexec      logs        README.txt  share
etc  lib      LICENSE.txt  NOTICE.txt  sbin

此时input和output不会出现在当前目录下,而是上传到了分布式文件系统中,网页上可以看到

[[email protected] hadoop]$ bin/hdfs dfs -cat output/*
1	dfsadmin
[[email protected] hadoop]$ bin/hdfs dfs -get output				##从分布式系统中get下来output目录
[[email protected] hadoop]$ cd output/
[[email protected] output]$ ls
part-r-00000  _SUCCESS
[[email protected] output]$ cat *
1	dfsadmin

三、分布式hadoop

1.首先停掉服务

[[email protected] hadoop]$ sbin/stop-dfs.sh			##我这里因为是关机以后第二天写的文章,所以不需要关闭该服务
[[email protected] ~]$ jps
2112 Jps
[[email protected] ~]$ cd /tmp/
[[email protected] tmp]$ rm -fr *								##删除所有的hadoop文件

2.新开hadoop2和hadoop3,新建用户

[[email protected] ~]# useradd -u 1000 hadoop
[[email protected] ~]# useradd -u 1000 hadoop

安装nfs-utils

[[email protected] ~]# yum install -y nfs-utils
[[email protected] ~]# yum install -y nfs-utils
[[email protected] ~]# yum install -y nfs-utils

开启服务

[[email protected] ~]# systemctl start rpcbind
[[email protected] ~]# systemctl start rpcbind
[[email protected] ~]# systemctl start rpcbind

3.hadoop1开启nfs存储,并共享

[[email protected] ~]# systemctl start nfs-server
[[email protected] ~]# vim /etc/exports
/home/hadoop   *(rw,anonuid=1000,anongid=1000)
[[email protected] ~]# exportfs -rv
exporting *:/home/hadoop
[[email protected] ~]# showmount -e
Export list for hadoop1:
/home/hadoop *

4.hadoop3、hadoop4挂载

[[email protected] ~]# mount 172.25.1.1:/home/hadoop /home/hadoop
[[email protected] ~]# df
Filesystem              1K-blocks    Used Available Use% Mounted on
/dev/mapper/rhel-root    17811456 1098404  16713052   7% /
devtmpfs                   497300       0    497300   0% /dev
tmpfs                      508264       0    508264   0% /dev/shm
tmpfs                      508264   13072    495192   3% /run
tmpfs                      508264       0    508264   0% /sys/fs/cgroup
/dev/sda1                 1038336  141476    896860  14% /boot
tmpfs                      101656       0    101656   0% /run/user/0
172.25.1.1:/home/hadoop  17811456 2796416  15015040  16% /home/hadoop

[[email protected] ~]# mount 172.25.1.1:/home/hadoop /home/hadoop
[[email protected] ~]# df
Filesystem              1K-blocks    Used Available Use% Mounted on
/dev/mapper/rhel-root    17811456 1097584  16713872   7% /
devtmpfs                   497300       0    497300   0% /dev
tmpfs                      508264       0    508264   0% /dev/shm
tmpfs                      508264   13072    495192   3% /run
tmpfs                      508264       0    508264   0% /sys/fs/cgroup
/dev/sda1                 1038336  141476    896860  14% /boot
tmpfs                      101656       0    101656   0% /run/user/0
172.25.1.1:/home/hadoop  17811456 2796416  15015040  16% /home/hadoop

这样我们就可以免密登陆hadoop2和hadoop3了,因为我们共享了家目录,里面含有ssh的公钥和私钥
5.重新编辑配置文件

[[email protected] ~]# su - hadoop
Last login: Mon Apr 15 23:12:09 EDT 2019 on pts/0
[[email protected] ~]$ cd hadoop/etc/hadoop/
[[email protected] hadoop]$ vim core-site.xml 
<configuration>
    <property>
        <name>fs.defaultFS</name>
        <value>hdfs://172.25.1.1:9000</value>
    </property>
</configuration>
[[email protected] hadoop]$ vim hdfs-site.xml 
<configuration>
    <property>
        <name>dfs.replication</name>
        <value>2</value>     ##改为两个节点
    </property>
</configuration>
[[email protected] hadoop]$ vim workers 
172.25.1.2
172.25.1.3

此时我们会发现其他节点的也有了这个文件,因为是共享的
hadoop分布式系统
6.格式化,并启动hadoop服务

[[email protected] hadoop]$ bin/hdfs namenode -format
[[email protected] hadoop]$ sbin/start-dfs.sh 
Starting namenodes on [hadoop1]
Starting datanodes
172.25.1.2: Warning: Permanently added '172.25.1.2' (ECDSA) to the list of known hosts.
172.25.1.3: Warning: Permanently added '172.25.1.3' (ECDSA) to the list of known hosts.
Starting secondary namenodes [hadoop1]
[[email protected] hadoop]$ jps
2902 SecondaryNameNode
2682 NameNode
3021 Jps

我们在hadoop2上也可以看到进程
hadoop分布式系统
7.测试

[[email protected] hadoop]$ bin/hdfs dfs -mkdir -p /user/hadoop
[[email protected] hadoop]$ bin/hdfs dfs -mkdir input
[[email protected] hadoop]$ bin/hdfs dfs -put etc/hadoop/*.xml input
[[email protected] hadoop]$ bin/hadoop jar share/hadoop/mapreduce/hadoop-mapreduce-examples-3.0.3.jar grep input output 'dfs[a-z.]+'

网页上查看,有两个数据节点
hadoop分布式系统
这里我的节点一直初步来,后来发现是我的解析有问题,我怕是感冒脑子坏了
hadoop分布式系统
hadoop4模拟客户端

[[email protected] hadoop]$ useradd -u  1000 hadoop
[[email protected] hadoop]$ yum  install -y  nfs-utils
[[email protected] hadoop]$ systemctl start rpcbind
[[email protected] hadoop]$ mount 172.25.14.1:/home/hadoop /home/hadoop
[[email protected] hadoop]$ su - hadoop
[[email protected] hadoop]$ vim /home/hadoop/hadoop/etc/hadoop/workers
172.25.1.2
172.25.1.3
172.25.1.4

然后在hadoop1把hadoop重启
hadoop分布式系统
新建并上传文件
hadoop分布式系统
然后在网页端查看

hadoop分布式系统
ok~