Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compute workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Engineering Computational Practices (Rev. 1)

This manual will attempt to motivate the use of an automated build system for the purposes of computational science and engineering. As part of this motivation, the surrounding computational practices of version control, documen tation, compute environment management, and regression testing will also be addressed as applied to the practice of computational engineering. Specifically, this manual intends to motivate the adoption of these traditional software engineering practices for use in research and production engineering simulation projects. This manual is not the first such effort in the greater scientific computing community. In fact, the authors relied heavily on the lesson plans of the Software Carpentry, established to teach computing skills to researchers in 1998. As the intention for this manual is to lay out fundamental practices of engineering computing, it will not attempt to fully teach the underlying concepts and will instead reference the well designed lesson plans of the Software Carpentry. Where possible, this manual will explain to general computing practices and concepts and limit discussion of specific software implementations to examples or vehicles for practice in concrete application. The specific software taught by the Software Carpentry curriculum is an excellent starting point to learn the core concepts of computational engi neering. However, the authors have found that applications to engineering simulation and analysis require translation of these software development concepts into the language and workflows of computational engineers. Adopting these computational tools may require engineers to re-imagine their workflows in some combination of traditional engineer ing and software concepts. It has also been necessary to extend existing software build systems for engineering practices beyond the simple wrap ping of engineering software execution. Where necessary, examples of specific software and their method of extension to engineering simulations will be given, with reference to the User Manual for recommended practical use. Where this manual relies on specific implementation examples, it should be understood that the practicing engineer may find that different software is more amenable to their specific work. It is always the overall collection of computational practices is more important than any specific software implementation. The ability to recognize which concepts are implemented by a software package will make a practicing engineer agile to changing project needs, computing resources, numeric solvers, programming languages, and even available funding.

42 ENGINEERING↗

2019 Computing Sciences Strategic Plan

Computing has transformed nearly every aspect of scientific inquiry — across disciplines and across scales — from the behavior of subatomic particles to the formation of structures in the early universe, from the assembly of the human genome to the evolution of earth systems. Over the past two decades, computing has become an integral part of how Berkeley Lab is “Bringing Science Solutions to the World.” Advances in computing and mathematics have been key, with new mathematical models of complex physical phenomena, new methods for analyzing complex data, new algorithms for accuracy and scaling and sophisticated software systems that encapsulate these techniques into open, reusable tools. The performance of NERSC computers and the ESnet network have grown by several orders of magnitude, along with our understanding of how to map scientific computations and workflows onto these systems. From research to facility operations, the passion, talent and dedication of the Computing Sciences Area staff has been the cornerstone of our success. The plan outlined in this document describes the next step in a journey to expand the influence and impact of our efforts, building an increasingly connected global enterprise for science that places more powerful instruments in the hands of scientists, along with more powerful methods and tools for modeling, analysis and prediction.

97 MATHEMATICS AND COMPUTING↗

What's Left for a Computational Chemist To Do in the Age of Machine Learning?

Machine learning (ML) has become a central focus of the computational chemistry community. In this paper, I will first discuss my personal history in the field. Then I will provide a broader view of how this resurgence in ML interest echoes and advances upon earlier efforts. Although numerous changes have brought about this latest wave, one of the most significant is the increased accuracy and efficiency of low-cost methods (e.g., density functional theory or DFT) that have made it possible to generate large data sets for ML models. ML has also been used to bypass, guide, or improve DFT. The field of computational chemistry thus finds itself at a crossroads as ML both augments and supersedes traditional efforts. I will present what I believe the role of the computational chemist will be in this evolving landscape, with specific focus on my experience in the development of autonomous workflows in computational materials discovery for open-shell transition-metal chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Geometric Modeler for Overlap Detection and Resolution (ASC IC L2 Milestone 7181 FY2020 Final Review)

The final review for the FY20 Advanced Simulation and Computing (ASC) Integrated Codes (IC) L2 Milestone #7181 was conducted on August 31, 2020 at Sandia National Laboratories in Albuquerque, New Mexico. The review panel unanimously agreed that the milestone has been successfully completed. Roshan Quadros (1543) led the milestone team and various members from the team presented the results. The review panel was comprised of staff from Sandia National Laboratories Albuquerque and California that are involved with computational engineering modeling and analysis. The panel consisted of experts in the fields of solid modeling, discretization, meshing, simulation workflows, and computational analysis including personnel Brett Clark (1543, Chair); Jay Foulk (8363); Jackie Moore (1553); Ron Kensek (1341); Ed Hoffman (8753); Dan Ibanez (1443). The presentation documented the technical approach of the team and summarized the results with sufficient detail to demonstrate both the value and the completion of the milestone. A separate SAND report was also generated with more detail to supplement the presentation. The purpose of the milestone was to advance capabilities for automatically finding, displaying, and resolving geometric overlaps in CAD models.

97 MATHEMATICS AND COMPUTING↗

WAVES [Slides]

WAVES (LANL code C23004) is a computational engineering workflow tool that integrates parametric studies with traditional software build systems.

42 ENGINEERING↗

Interactive Supercomputing With Jupyter

Rich user interfaces like Jupyter have the potential to make interacting with a supercomputer easier and more productive, consequently attracting new kinds of users and helping to expand the application of supercomputing to new science domains. For the scientist-user, the ideal rich user interface delivers a familiar, responsive, introspective, modular, and customizable platform upon which to build, run, capture, document, re-run, and share analysis workflows. From the provider or system administrator perspective, such a platform would also be easy to configure, deploy securely, update, customize, and support. Jupyter checks most if not all of these boxes. But from the perspective of leadership computing organizations that provide supercomputing power to users, such a platform should also make the unique features of a supercomputer center more accessible to users and more composable with high performance computing (HPC) workflows. Project Jupyter’s core design philosophy of extensibility, abstraction, and agnostic deployment, has allowed HPC centers like NERSC to bring in advanced supercomputing capabilities that can extend the interactive notebook environment. This has enabled a rich scientific discovery platform, particularly for experimental facility data analysis and machine learning problems.

97 MATHEMATICS AND COMPUTING↗

Dial

A key step in almost all scientific endeavors is answering the question: Given this data I already collected, what new data do I expect will yield the most useful information toward my scientific objective? The area of (sequential) experimental design has long been investigating answers to this question, but in recent years techniques from the machine learning subfield of active learning are increasingly applied. Researchers need a simple software tool for active learning applied to experimental design that can easily integrate into their existing workflows. This computer code, Dial, provides a microservice in ORNL's INTERSECT ecosystem for active learning applied to experimental design. By being part of the INTERSECT ecosystem, Dial is simple to integrate into any INTERSECT-based workflow. Dial provides multiple backend options, where a backend is an implementation of a specific active learning method. Users can select the backend that performs best for their application. Developers can also add new backends as needed. At its core, Dial receives a set of pre-existing measurements and input parameter bounds and then recommends one or more new sets of parameters to measure. Dial also includes interfaces to other microservices in the INTERSECT ecosystem so that it can be incorporated into INTERSECT campaigns. Dial provides a simple, yet powerful interface to convert automated INTERSECT workflows into autonomous workflows that adapt based on the results that are obtained. A shared microservice for active learning prevents duplicated effort by each application team implementing its own adaptive design of experiments tool.

Drane, Lance [Oak Ridge National Laboratory (ORNL)↗

Workflows Community Summit 2022: A Roadmap Revolution

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex scientific experiments that range from the execution of a cloud-based data preprocessing pipeline to multi-facility instrument-to-edge-to-HPC computational workflows. Given the changing landscape of scientific computing (often referred to as a computing continuum) and the evolving needs of emerging scientific applications, it is paramount that the development of novel scientific workflows and system functionalities seek to increase the efficiency, resilience, and pervasiveness of existing systems and applications. Specifically, the proliferation of machine learning/artificial intelligence (ML/AI) workflows, need for processing large-scale datasets produced by instruments at the edge, intensification of near real-time data processing, support for long-term experiment campaigns, and emergence of quantum computing as an adjunct to HPC, have significantly changed the functional and operational requirements of workflow systems. Workflow systems now need to, for example, support data streams from the edge-to-cloud-to-HPC, enable the management of many small-sized files, allow data reduction while ensuring high accuracy, orchestrate distributed services (workflows, instruments, data movement, provenance, publication, etc.) across computing and user facilities, among others. Further, to accelerate science, it is also necessary that these systems implement specifications/standards and APIs for seamless (horizontal and vertical) integration between systems and applications, as well as enable the publication of workflows and their associated products according to the FAIR principles.

97 MATHEMATICS AND COMPUTING↗

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS↗

Asynchronous Execution of Heterogeneous Tasks in ML-Driven HPC Workflows

Heterogeneous scientific workflows consist of numerous types of tasks that require execution on heterogeneous resources. Asynchronous execution of those tasks is crucial to improve resource utilization, task throughput and reduce workflows' makespan. Therefore, middleware capable of scheduling and executing different task types across heterogeneous resources must enable asynchronous execution of tasks. In this paper, we investigate the requirements and properties of the asynchronous task execution of machine learning (ML)-driven high-performance computing (HPC) workflows. We model the degree of asynchronicity permitted for arbitrary workflows and propose key metrics that can be used to determine qualitative benefits when employing asynchronous execution. Our experiments represent relevant scientific drivers, we perform them at scale on Summit, and we show that the performance enhancements due to asynchronous execution are consistent with our model.

97 MATHEMATICS AND COMPUTING↗

Scaling SQL to the Supercomputer for Interactive Analysis of Simulation Data

AI and simulation workloads consume and generate large amounts of data that need to be searched, transformed and merged with other data. With the goal of treating data as a first-class citizen inside a traditionally compute-centric HPC environment, we explore how the use of accelerators and high-speed interconnects can speed up tasks which otherwise constitute bottlenecks in computational discovery workflows. BlazingSQL is SQL engine that runs natively on NVIDIA GPUs and supports internode communication for fast analytics on terabyte-scale tabular data sets. We show how a fast interconnect improves query performance if leveraged through the Unified Communication X (UCX) middleware. We envision that future computing platforms will integrate accelerated database query capabilities for immediate and interactive analysis of large simulation data.

Glaser, Jens↗

Machine learned potential for high-throughput phonon calculations of metal—organic frameworks

Metal–organic frameworks (MOFs) are highly porous and versatile materials studied extensively for applications such as carbon capture and water harvesting. However, computing phonon-mediated properties in MOFs, like thermal expansion and mechanical stability, remains challenging due to the large number of atoms per unit cell, making traditional Density Functional Theory (DFT) methods impractical for high-throughput screening. Recent advances in machine learning potentials have led to foundation atomistic models, such as MACE-MP-0, that accurately predict equilibrium structures but struggle with phonon properties of MOFs. In this work, we developed a workflow for computing phonons in MOFs within the quasi-harmonic approximation with a fine-tuned MACE model, MACE-MP-MOF0. The model was trained on a curated dataset of 127 representative and diverse MOFs. The fine-tuned MACE-MP-MOF0 improves the accuracy of phonon density of states and corrects the imaginary phonon modes of MACE-MP-0, enabling high-throughput phonon calculations with state-of-the-art precision. The model successfully predicts thermal expansion and bulk moduli in agreement with DFT and experimental data for several well-known MOFs. These results highlight the potential of MACE-MP-MOF0 in guiding MOF design for applications in energy storage and thermoelectrics.

Elena, Alin Marin↗

Illuminating Ligand Field Contributions to Molecular Qubit Spin Relaxation via T 1 Anisotropy

Electron spin relaxation in paramagnetic transition metal complexes constitutes a key limitation on the growth of molecular quantum information science. However, there exist very few experimental observables for probing spin relaxation mechanisms, leading to a proliferation of inconsistent theoretical models. Here we demonstrate that spin relaxation anisotropy in pulsed electron paramagnetic resonance is a powerful spectroscopic probe for molecular spin dynamics across a library of highly coherent Cu(II) and V(IV) complexes. Here, neither the static spin Hamiltonian anisotropy nor contemporary computational models of spin relaxation are able to account for the experimental T 1 anisotropy. Through analysis of the spin-orbit coupled wavefunctions, we derive an analytical theory for the T 1 anisotropy that accurately reproduces the average experimental anisotropy of 2.5. Furthermore, compound-by-compound deviations from the average anisotropy provide a promising approach for observing specific ligand field and vibronic excited state coupling effects on spin relaxation. Finally, we present a simple density functional theory workflow for computationally predicting T 1 anisotropy. Analysis of spin relaxation anisotropy leads to deeper fundamental understanding of spin-phonon coupling and relaxation mechanisms, promising to complement temperature-dependent relaxation rates as a key metric for understanding molecular spin qubits.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Netload Range Cost Curves for Coordinated Transmission-Distribution Planning Under DER Growth Uncertainty

The increasing penetration of distributed energy resources (DERs) requires better coordination between transmission and distribution (T&D) planning to ensure system security and cost efficiency. However, misaligned planning horizons, computational burdens, and privacy concerns hinder effective coordination, leading to either underutilized resources caused by overinvestments or reliability risks due to underinvestment. To address this challenge, we introduce netload range cost curves (NRCCs), a novel approach for managing long-term DER growth uncertainty through T&D coordination, while preserving existing data-sharing and regulatory structures. NRCCs provide pairs of (i) peak substation netload guarantees and (ii) corresponding distribution upgrade options and costs, enabling their seamless integration into transmission planning workflows. To compute NRCCs efficiently, we develop a transmission-aware distribution network planning (TADNP), which is subsequently integrated to an iterative computation procedure. These NRCCs are then embedded into an NRCC-informed transmission planning model to enable resource-efficient coordination. We illustrate our proposed approach with a case study based on realistic distribution and transmission systems in the San Francisco Bay Area, California. Our results indicate the possibility of dramatic savings in transmission investments by incorporating the proposed NRCC-integrated T&D coordination framework.

Li, Yujia↗

ExaWorks software development kit: a robust and scalable collection of interoperable workflows technologies

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different types of tasks (e.g., simulation, analysis, and learning) that need to be mapped, scheduled, and launched on different computing. That requires a software stack that enables users to code their workflows and automate resource management and workflow execution. Currently, there are many workflow technologies with diverse levels of robustness and capabilities, and users face difficult choices of software that can effectively and efficiently support their use cases on HPC machines, especially when considering the latest exascale platforms. We contributed to addressing this issue by developing the ExaWorks Software Development Kit (SDK). The SDK is a curated collection of workflow technologies engineered following current best practices and specifically designed to work on HPC platforms. We present our experience with (1) curating those technologies, (2) integrating them to provide users with new capabilities, (3) developing a continuous integration platform to test the SDK on DOE HPC platforms, (4) designing a dashboard to publish the results of those tests, and (5) devising an innovative documentation platform to help users to use those technologies. Our experience details the requirements and the best practices needed to curate workflow technologies, and it also serves as a blueprint for the capabilities and services that DOE will have to offer to support a variety of scientific heterogeneous workflows on the newly available exascale HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗

Multi-scale computational screening and mechanistic insights of cyclic amines as solvents for improved lignocellulosic biomass processing

A computational screening workflow for the efficient deconstruction of cellulose, lignin and hemicellulose fractions of lignocellulosic biomass using cyclic amines as solvents. Lignocellulosic biomass is a promising feedstock for production of affordable fuels and chemicals from renewable resources. Effective solubilization and subsequent deconstruction of its cellulose, hemicellulose, and lignin fractions is essential for the viability of future biorefineries. This study used quantum chemistry-based equilibrium thermodynamics methods to evaluate the potential of 650 cyclic amines to solubilize cellulose, hemicellulose, and lignin. The activity coefficients of solvent - biopolymer interactions were predicted using the COSMO-RS (COnductor-like Screening MOdel for Real Solvents) method and used to identify cyclic amines that can efficiently dissolve and extract selective fractions of biopolymers during biomass pretreatment. Among the 650 cyclic amines, 1-piperazineethanmaine was predicted to be an effective solvent for extracting all three polymers and was experimentally shown to achieve the highest lignin removal (97.1%). Non-covalent interaction, reduced density gradient and quantum chemical calculations were performed to elucidate the dissolution mechanism of lignin, cellulose and hemicellulose and gain further molecular level insights into the interactions between the cyclic amines and biomass polymers that promote efficient solubilization and extraction. These analyses indicated that 1-piperazineethanmaine and 1-methylimidazole make noncovalent van der Waals, electrostatic interactions and hydrogen bonding with lignin, leading to enhanced lignin removal, while the strong intramolecular hydrogen bonding interactions in cellulose and hemicellulose result in weaker solvent-biopolymer interactions. Overall, the computational approach provided an efficient method for identifying cyclic amines tailored for optimal biomass pretreatment and resulted in the identification of a potential new class of solvents for effective biomass pretreatment.

Kumar, Nikhil↗

ExaWorks: Workflows for Exascale

Exascale computers will offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discovery and insight. These software combinations and integrations, however, are difficult to achieve due to challenges of coordination and deployment of heterogeneous software components on diverse and massive platforms. We present the ExaWorks project, which can address many of these challenges: ExaWorks is leading a co-design process to create a workflow Software Development Toolkit (SDK) consisting of a wide range of workflow management tools that can be composed and interoperate through common interfaces. We describe the initial set of tools and interfaces supported by the SDK, efforts to make them easier to apply to complex science challenges, and examples of their application to exemplar cases. Furthermore, we discuss how our project is working with the workflows community, large computing facilities as well as HPC platform vendors to sustainably address the requirements of workflows at the exascale.

97 MATHEMATICS AND COMPUTING↗