Unified batch and streaming: traditional vs. Deephaven
Unified API eliminates complexity of separate batch and real-time systems
Traditional: Complex Multi-System
Batch Layer
• Historical data
• Parquet, HDFS
• Spark/Hadoop
• Hours/days latency
Speed Layer
• Real-time streams
• Kafka, Flink
• Storm, Samza
• ms-second latency
merges into
Serving Layer
• Merges batch + real-time views
• Complex data reconciliation
• Different APIs for each layer
• Manual consistency management
Challenges:
• Multiple systems to learn & maintain
• Different APIs for batch vs streaming
• Complex data reconciliation logic
• Infrastructure overhead
Deephaven: Unified Single System
Deephaven Query Engine
Batch Data
• Parquet, CSV
• Historical tables
• Static sources
Real-time Data
• Kafka streams
• Live tables
• Ticking sources
Unified DAG-based Update Model
Same API • Same semantics • Automatic consistency
Single Table API
.where() .agg_by() .join() - works on all data
Benefits:
• One system, one API to learn
• Seamless batch + real-time integration
• Automatic consistency via DAG
• Lower operational complexity
Code example: same API for both
Python code works identically for batch and real-time data:
Batch (Historical):
historical = read_csv("trades_2024.csv")
result = historical.where("Price > 100")
.agg_by([agg.avg("Price")], by=["Symbol"])
Real-time (Live):
live = consume_kafka(kafka_config, "trades",
key_spec=KeyValueSpec.IGNORE, value_spec=trade_spec,
table_type=TableType.append())
result = live.where("Price > 100")
.agg_by([agg.avg("Price")], by=["Symbol"])
Identical operations! The only difference is the data source.
Both use the same DAG, both maintain consistency, both support the same table operations.