hive整合hbase

Hive整合HBase后的使用场景:

通过Hive把数据加载到HBase中,数据源可以是文件也可以是Hive中的表。

通过整合,让HBase支持JOIN、GROUP等SQL查询语法。

通过整合,不仅可完成HBase的数据实时查询,也可以使用Hive查询HBase中的数据完成复杂的数据分析。

配置

因为Hive与HBase整合的实现是利用两者本身对外的API接口互相通信来完成的,其具体工作交由Hive的lib目录中的hive-hbase-handler-.jar工具类来实现。所以只需要将hive的 hive-hbase-handler-.jar 复制到hbase/lib中就可以了。

 [root@host lib]# cp hive-hbase-handler-2.1.1.jar $HBASE_HOME/lib

测试

通过hive创建hbase表

hive> CREATE TABLE t_name (id INT, NAME string)
    >      stored BY 'org.apache.hadoop.hive.hbase.HBaseStorageHandler'
    >     WITH serdeproperties (
    >     "hbase.columns.mapping" = ":key,st1:name")
    >    tblproperties ("hbase.table.name" = "t_name","hbase.mapred.output.outputtable" = "t_name");
OK
Time taken: 1.625 seconds

在hive中查看:

hive> show tables;
OK
cust_copy
t_name
Time taken: 0.127 seconds, Fetched: 2 row(s)

hive> show create table t_name;
OK
CREATE TABLE `t_name`(
  `id` int COMMENT '',
  `name` string COMMENT '')
ROW FORMAT SERDE
  'org.apache.hadoop.hive.hbase.HBaseSerDe'
STORED BY
  'org.apache.hadoop.hive.hbase.HBaseStorageHandler'
WITH SERDEPROPERTIES (
  'hbase.columns.mapping'=':key,st1:name',
  'serialization.format'='1')
TBLPROPERTIES (
  'COLUMN_STATS_ACCURATE'='{\"BASIC_STATS\":\"true\"}',
  'hbase.mapred.output.outputtable'='t_name',
  'hbase.table.name'='t_name',
  'numFiles'='0',
  'numRows'='0',
  'rawDataSize'='0',
  'totalSize'='0',
  'transient_lastDdlTime'='1526546542')
Time taken: 0.308 seconds, Fetched: 19 row(s)

在HBASE中查看

hbase(main):004:0> list 't_name'
TABLE                                                                                                                                        
t_name                                                                                                                                       
1 row(s)
Took 0.0092 seconds                                                                                                                          
=> ["t_name"]

在hbase插入数据并查看数据:

hbase(main):006:0> put 't_name','1','st1:name','xiaoma'
Took 0.3709 seconds                                                                                                                          
hbase(main):007:0> put 't_name','2','st1:name','xiaozhang'
Took 0.0038 seconds                                                                                                                          
hbase(main):008:0> put 't_name','3','st1:name','tianyongtao'
Took 0.0051 seconds

hbase(main):009:0> scan 't_name'
ROW                                  COLUMN+CELL                                                                                             
 1                                   column=st1:name, timestamp=1526547097913, value=xiaoma                                                  
 2                                   column=st1:name, timestamp=1526547115702, value=xiaozhang                                               
 3                                   column=st1:name, timestamp=1526547130241, value=tianyongtao                                             
3 row(s)
Took 0.0327 seconds 

通过hive查询:

hive> select * from t_name;
OK
t_name.id       t_name.name
1       xiaoma
2       xiaozhang
3       tianyongtao
Time taken: 0.414 seconds, Fetched: 3 row(s)
hive> select * from t_name where id=1;
OK
t_name.id       t_name.name
1       xiaoma
Time taken: 1.246 seconds, Fetched: 1 row(s)
hive> select * from t_name where id>1;
OK
t_name.id       t_name.name
2       xiaozhang
3       tianyongtao
Time taken: 0.383 seconds, Fetched: 2 row(s)

删除表测试:

hive> drop table t_name;
OK
Time taken: 1.851 seconds

经查hbase中的t_name表被同步删除了

猜你喜欢

转载自www.cnblogs.com/playforever/p/9051796.html
今日推荐