Carter Shanklin gives us a few patterns for updating tables in Hive:
Historically, keeping data up-to-date in Apache Hive required custom application development that is complex, non-performant, and difficult to maintain. HDP 2.6 radically simplifies data maintenance with the introduction of SQL MERGE in Hive, complementing existing INSERT, UPDATE, and DELETE capabilities.
This article shows how to solve common data management problems, including:
-
Hive upserts, to synchronize Hive data with a source RDBMS.
-
Update the partition where data lives in Hive.
-
Selectively mask or purge data in Hive.
This isn’t the Hive of 2013; it’s much closer to a real-time warehouse.