Heterogeneous Stream Processing Engine within the PolyDBMS Architecture (Master Thesis, Ongoing)

Author

Mathieu Groenen

Description

With the increasing digitalisation, from sensors gathering real-time meteorological data to businesses gathering real-time data about orders, shipping and the like, a strong desire exists to have systems capable of efficiently and correctly handling fast moving, volatile and unbounded data, i.e., data theoretically infinite in size. Systems, called Stream Processing (SP) systems, provide such capabilities. They provide near real-time processing of fast moving data and guarantees like exactly-once processing, event-time ordering with replication and fault tolerance. SP systems are also capable of a wide variety of operations, from mathematical to conditional, with operators which can be stateless or stateful and capable of

distribution and load balancing. While on the surface these guarantees are nice, substantial engineering effort is still required to build reliable and correct data pipelines and even more effort considering the migration of existing pipelines to such SP systems. An already existing, static data pipeline, would need a way to transform static data into streaming data, making it ingestible for SP systems. Additionally, using a SP systems would create issues and semantics that weren’t there before, for example the issue of being unable to gather all data due to the streams unbounded nature and following from this, guaranteeing that all data in a specified time frame has arrived, and that no data will arrive late. These challenges stand in contrast to Database Management Systems (DBMS), already used widely in existing data pipelines as a storage engine while providing correctness and reliability together with sufficient speeds for different workloads.

Thus, the question arises as to why not just adapt the SP system to fit the data pipeline, instead of adapting the existing static data pipeline to fit the SP systems streaming model. If DBMS are already widely used in existing data pipelines and provide proper storage as well as fault tolerance and replication guarantees with strong read and write speeds, why not use the DBMS of the already existing pipelines as both the storage and processing engine simultaneously. The DBMS would then take over the role of performing the data processing tasks of a SP system, which it could do directly on the data it has stored, instead of having to transform the static data into streaming data for it to be ingestible. The objective of this thesis is to design such a SP system, a database integrated stream processing system.

Start / End Dates

2026/10/12 - 2027/04/11

Supervisors

Research Topics