Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data science workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗

Using Selection Pressure as an Asset to Develop Reusable, Adaptable Software Systems

The Goddard Earth Sciences Data and Information Services Center (GES DISC) at NASA has over the years developed and honed several reusable architectural components for supporting large-scale data centers with a large customer base. These include a processing system (S4PM) and an archive system (S4PA) based upon a workflow engine called the Simple Scalable Script based Science Processor (S4P) and an online data visualization and analysis system (Giovanni). These subsystems are currently reused internally in a variety of combinations to implement customized data management on behalf of instrument science teams and other science investigators. Some of these subsystems (S4P and S4PM) have also been reused by other data centers for operational science processing. Our experience has been that development and utilization of robust interoperable and reusable software systems can actually flourish in environments defined by heterogeneous commodity hardware systems the emphasis on value-added customer service and the continual goal for achieving higher cost efficiencies. The repeated internal reuse that is fostered by such an environment encourages and even forces changes to the software that make it more reusable and adaptable. Allowing and even encouraging such selective pressures to software development has been a key factor In the success of S4P and S4PM which are now available to the open source community under the NASA Open source Agreement

Berrick, Stephen↗

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

Preparation of the Multi-Site Data Processing at the Vera C. Rubin Observatory

The Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) Camera is scheduled to start taking data in the summer of 2025. The Data Release Production will run the LSST Science Pipe software at data facilities in the US, France and the UK. The LSST Science Pipeline consists of complex directed acyclic graphs (DAGs) of tasks. Rubin will use the Production and Distributed Analysis (PanDA) workflow and workload management system to orchestrate this complex workflow and the distribution of workloads to the data facilities. When run end-to-end by a team of data production staff, this processing (the Science Pipelines, distributed by the workflow and workload management system) is referred to as a 'campaign'. This paper describes the central services and data facility specific services that support this multi-site data process model, including the service deployment infrastructure, the workload and workflow system, the Campaign Management tools, and connection to Rubin Data Management. This paper will also mention the experience of processing the Rubin Commissioning Camera data. All these are part of the effort to scale up the processing capabilities for the expected very large data volume from the LSST Camera.

Yang, Wei [SLAC]↗

Future Trends in Nuclear Physics Computing

In nuclear physics (NP) today the study of quarks, gluons and their strong interactions extends across a broad research program at a varied range of collaborative scales, from a few collaborators up to large experiments at scales comparable to those typical of high energy physics (HEP). Overall, the software and computing efforts vary accordingly, from pragmatic do-it-yourself approaches among a few, to substantial organized software and computing activities within large experiments. With new experiments starting up and on the horizon [1], and rapidly increasing data volumes [2, 3] and processing demands even at small experiments, the NP community has in recent years been thinking about the next generation of data processing and analysis workflows that will maximize the science output. One context for this discussion has been a series of workshops, “Future Trends in Nuclear Physics Computing” [4]. The most recent in this series took place in Fall 2020, organized by the authors together with colleagues. The workshop focused on identifying the unique aspects of software and computing in NP, and discussing how the NP community could strengthen common efforts and chart a path forward for the next decade, sure to be an exciting one with rich ongoing scientific programs at Brookhaven National Laboratory (BNL), Jefferson Lab (JLab), and other NP facilities, and culminating in datataking at the Electron-Ion Collider (EIC) [5,6,7] in the early 2030s. Without claiming to present a collective view from the workshop and discussions since—fortunately this is not expected of us in this opinion editorial—we offer here our reflections on the topic, informed by the workshop and the summary we authored with our colleagues [8], as well as discussions and developments in the eventful time since.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Intercomparing Open Source Surface Water Extent Mapping Products & Software Packages

The open source revolution of Earth Observation (EO) science has resulted in increased openness of EO data, workflows to transform that data into end products (e.g. surface water extent maps), and the products themselves. However, this revolution has oversaturated decision-makers with products that can give conflicting results. Thus, it is increasingly crucial for scientists to communicate their methodologies and assumptions so scientific products can be used accurately. Recognizing this challenge, SERVIR – a joint initiative between NASA, USAID, and geospatial organizations in Asia, Africa, and Latin America – is conducting a regional intercomparison of open source surface water extent products and packages. SERVIR’s Hindu Kush Himalaya and Southeast Asia “hubs” have developed satellite-based surface water mapping services involving customizable code packages that are operationally run at each hub. These services are regionally and locally tailored to inform specific decisions and early actions. Conversely, the scientific community has released surface water products that are global or near-global, but are not customizable. These packages and products employ different methodologies and sensors, causing decision-makers to evaluate trade-offs related to physical sensor characteristics (e.g. spectral, temporal, and spatial resolution, and latency). We will discuss the tradeoffs, strengths, and weaknesses of the sensor characteristics and methodologies associated with each product/package, and provide preliminary results of a validation effort intercomparing products/packages for case studies in South and Southeast Asia. Understanding the strengths and weaknesses of these products is crucial in both the aftermath of a flood event and in preparing for future floods.

Micky Maganini↗

Management and Storage of Scientific Data

Scientific discoveries rely heavily on efficient access, search, and management of massive data sets. Data management technologies have, for decades, provided foundational capabilities for scientific computing. Just as storage, input/output (I/O), and data management have been fundamental to simulation-based science for many years, so too are capable data-management technologies key to the success of today’s scientific workflows utilizing data intensive and machine learning (ML) techniques. The Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program has invested broadly in data-management research focused on high-performance computing (HPC) systems, from parallel file systems that store data to application software that makes these systems more productive. Still, advances in technology combined with growing diversity of supported science strongly motivate continued investment in this area. In January 2022, ASCR convened a workshop to identify priority research directions in the area of data management for high-performance and scientific computing. Attendees were challenged to identify promising approaches that would support the breadth of the DOE mission, including the explosion of artificial intelligence (AI) uses and the growing needs of experimental and observational science. Technological and science drivers were identified and considered as they relate to key aspects of data management such as interfaces, architectural design, and FAIR principles (Findable, Accessible, Interoperable, and Reusable). The thoughts of the workshop participants were distilled into a set of four priority research directions with the potential for high impact on DOE science. These research directions are summarized in the following pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

ESnet-JLab FPGA Accelerated Transport (data plane) [EJFAT (udplb)] v1.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplb) implements the data plane portion of the EJFAT system. It is an FPGA design that rewrites and forwards data packets from a UDP-based scientific workflow to high-performance compute nodes. It depends on another program (udplbd, disclosed separately) to implement the control system.

Bengough, Peter [Malleable Networks, Inc.]↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Toward a Seamless Integration of Computing, Experimental, and Observational Science Facilities: A Blueprint to Accelerate Discovery

The Department of Energy, Office of Science operates world-leading facilities for experimental, observational, and computational science. DOE supercomputing facilities will reach performance at the scale of ExaFLOPs in the coming years, enabling new vistas of scale and precision for large scale simulations and data analysis. Experimental scientific facilities are undergoing similar upgrades that will lead to higher data rates and correspondingly larger computational demands, and will increase the need for near-real-time processing and resilient support for more complex workflows. A transformation of science is underway, with workloads at supercomputing facilities increasingly driven by this explosion of data from instruments and experimental facilities, as well as the accelerating use of Artificial Intelligence (AI) as a tool for scientific discovery. A seamless integration of computing, networking, instruments, and experimental facilities is required to support these emerging workloads and open up a new frontier of U.S. leadership in scientific discovery. We propose to accomplish this by providing frictionless access to the ASCR supercomputing facilities. We describe our vision of combining the power of ASCR supercomputers and networking infrastructure into an integrated scalable fabric, available to end user scientists via interfaces that aim to automate and simplify access to high performance computing systems. This will enable unprecedented computational science capabilities for experimental and observational facilities, and will create new opportunities to combine large simulations and modeling with experimental facility data analysis. This blueprint for creating an integrated network of computational and experimental facilities will provide an enriched discovery environment and open doors for new scientific communities to access the DOE’s world-leading computing and networking capabilities.

97 MATHEMATICS AND COMPUTING↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence↗

Discovery of complex oxides via automated experiments and data science

Significance Automation is accelerating the discovery of useful materials, yet testing even a small fraction of the billions of possible materials for a desired property is beyond the reach of workflows involving resource-intensive property measurements. Due to relationships among composition, structure, and properties, identifying a complex material with one interesting property makes it the proverbial needle in a haystack that merits testing for additional properties. We accelerate materials synthesis and optical characterization by employing physics-aware data science to identify materials for further investigation. With this approach, one does not need high-throughput methods for measuring every material property of interest since a single ultra-high–throughput workflow can guide material selection for other properties, which is a new paradigm for accelerated materials discovery.

36 MATERIALS SCIENCE↗

Distinguishing Provenance Equivalence of Earth Science Data

Reproducibility of scientific research relies on accurate and precise citation of data and the provenance of that data. Earth science data are often the result of applying complex data transformation and analysis workflows to vast quantities of data. Provenance information of data processing is used for a variety of purposes, including understanding the process and auditing as well as reproducibility. Certain provenance information is essential for producing scientifically equivalent data. Capturing and representing that provenance information and assigning identifiers suitable for precisely distinguishing data granules and datasets is needed for accurate comparisons. This paper discusses scientific equivalence and essential provenance for scientific reproducibility. We use the example of an operational earth science data processing system to illustrate the application of the technique of cascading digital signatures or hash chains to precisely identify sets of granules and as provenance equivalence identifiers to distinguish data made in an an equivalent manner.

Tilmes, Curt↗

Data Science Meets Physical Organic Chemistry

At the heart of synthetic chemistry is the holy grail of predictable catalyst design. In particular, researchers involved in reaction development in asymmetric catalysis have pursued a variety of strategies toward this goal. This is driven by both the pragmatic need to achieve high selectivities and the inability to readily identify why a certain catalyst is effective for a given reaction. While empiricism and intuition have dominated the field of asymmetric catalysis since its inception, enantioselectivity offers a mechanistically rich platform to interrogate catalyst-structure response patterns that explain the performance of a particular catalyst or substrate. In the early stages of an asymmetric reaction development campaign, the overarching mechanism of the reaction, catalyst speciation, the turnover limiting step, and many other details are unknown or posited based on related reactions. Considering the unclear details leading to a successful reaction, initial enantioselectivity data are often used to intuitively guide the ultimate direction of optimization. However, if the conditions of the Curtin-Hammett principle are satisfied, then measured enantioselectivity can be directly connected to the ensemble of diastereomeric transition states (TSs) that lead to the enantiomeric products, and the associated free energy difference between competing TSs (ΔΔ G ‡ = - RT ln[( S )/( R )], where ( S ) and ( R ) represent the concentrations of the enantiomeric products). We, and others, speculated that this important piece of information can be leveraged to guide reaction optimization in a quantitative way. Although traditional linear free energy relationships (LFERs), such as Hammett plots, have been used to illuminate important mechanistic features, we sought to develop data science derived tools to expand the power of LFERs in order to describe complex reactions frequently encountered in modern asymmetric catalysis. Specifically, we investigated whether enantioselectivity data from a reaction can be quantitatively connected to the attributes of reaction components, such as catalyst and substrate structural features, to harness data for asymmetric catalyst design. In this context, we developed a workflow to relate computationally derived features of reaction components to enantioselectivity using data science tools. The mathematical representation of molecules can incorporate many aspects of a transformation, such as molecular features from substrate, product, catalyst, and proposed transition states. Statistical models relating these features to reaction outputs can be used for various tasks, such as performance prediction of untested molecules. Perhaps most importantly, statistical models can guide the generation of mechanistic hypotheses that are embedded within complex patterns of reaction responses. Overall, merging traditional physical organic experiments with statistical modeling techniques creates a feedback loop that enables both evaluation of multiple mechanistic hypotheses and future catalyst design. In this Account, we highlight the evolution and application of this approach in the context of a collaborative program based on chiral phosphoric acid catalysts (CPAs) in asymmetric catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Challenges for Implementing FAIR Digital Objects with High Performance Workflows

New types of workflows are being used in science that couple traditional distributed and high-performance computing (HPC) with data-intensive approaches, and orchestrate ensembles of numerical simulations and artificial intelligence (AI) models. Such workflows may use AI models to supplement computation where numerical simulations may be too computationally expensive, to automate trivial yet time consuming operations, to perform preliminary selections among intractable numbers of combinations in domains as diverse as protein binding, fine-grid climate simulations, and drug discovery.

97 MATHEMATICS AND COMPUTING↗