DAPHNE - Integrated Data Analysis Pipelines for Large-Scale Data Management, HPC, and Machine Learning

Project Coordinator

Know-Center GmbH (Austria)

Description

Modern data-driven applications leverage large, heterogeneous data collections to find interesting patterns, and build robust machine learning (ML) models for accurate predictions. Large data sizes and advanced analytics spurred the development and adoption of data-parallel computation frameworks like Apache Spark or Flink as well as distributed ML systems like MLlib, TensorFlow, or PyTorch.

A key observation is that these new systems share many techniques with traditional high-performance computing (HPC), and the architecture of underlying HW clusters converges. Yet, the programming paradigms, cluster resource management, as well as data formats and representations differ substantially across data management, HPC, and ML software stacks. There is a trend though, toward complex data analysis pipelines that combine these different systems. Examples are workflows of distributed data pre-processing, tuned HPC libraries, and dedicated ML systems, but also HPC applications that leverage ML models for more cost-effective simulation. Major obstacles are:

limited development productivity for integrated analysis pipelines due to different programming models, and separated cluster environments
unnecessary data movement overhead and underutilization due to separate, statically provisioned clusters,
lack of a common system infrastructure with good interoperability.

For these reasons, DAPHNE’s overall objective is the definition of an open and extensible systems infrastructure for integrated data analysis pipelines. We aim at building a reference implementation of language abstractions (i.e., APIs and a domain-specific language), an intermediate representation, as well as compilation and runtime techniques with support for integrating and scheduling heterogeneous accelerator and storage devices.

A variety of real-world, high-impact use cases, datasets, and a new benchmark will be used for qualitative and quantitative analysis compared to state-of-the-art.

Cookie	Type	Duration	Description
_ga	third party	2 years	This cookie is installed by Google Analytics. The cookie is used to calculate visitor, session, camapign data and keep track of site usage for the site's analytics report. The cookies store information anonymously and assigns a randoly generated number to identify unique visitors.
_gat	persistent	1 year	Google uses this cookie to distinguish users.
_gid	third party	1 day	This cookie is installed by Google Analytics. The cookie is used to store information of how visitors use a website and helps in creating an analytics report of how the wbsite is doing. The data collected including the number visitors, the source where they have come from, and the pages viisted in an anonymous form.
cookielawinfo-checkbox-non-necessary		1 year	The cookie checks if the visitor is logged in or not. The purpose of the verification is to ensure security and prevent unauthorized changes to the site.
pll_language		1 year	Currently used website language.
wfwaf-authcookie-*		1 day	The cookie checks if the visitor is logged in or not. The purpose of the verification is to ensure security and prevent unauthorized changes to the site.
wp-settings		1 year	The website stores a cookie with an individual user ID from the user database table. This is used to customize the view of the admin interface and possibly the main interface of the site.

DAPHNE – Integrated Data Analysis Pipelines for Large-Scale Data Management, HPC, and Machine Learning

Researchers

More information

Follow us

Get in touch

Sign up for newsletter