Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compute workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

De novo design and Rosetta-based assessment of high-affinity antibody variable regions (Fv) against the SARS-CoV -2 spike receptor binding domain ( RBD )

The continued emergence of new SARS-CoV-2 variants has accentuated the growing need for fast and reliable methods for the design of potentially neutralizing antibodies (Abs) to counter immune evasion by the virus. Here, we report on the de novo computational design of high-affinity Ab variable regions (Fv) through the recombination of VDJ genes targeting the most solvent-exposed hACE2-binding residues of the SARS-CoV-2 spike receptor binding domain (RBD) protein using the software tool OptMAVEn-2.0. Subsequently, we carried out computational affinity maturation of the designed variable regions through amino acid substitutions for improved binding with the target epitope. Immunogenicity of designs was restricted by preferring designs that match sequences from a 9-mer library of “human Abs” based on a human string content score. We generated 106 different antibody designs and reported in detail on the top five that trade-off the greatest computational binding affinity for the RBD with human string content scores. We further describe computational evaluation of the top five designs produced by OptMAVEn-2.0 using a Rosetta-based approach. We used Rosetta SnugDock for local docking of the designs to evaluate their potential to bind the spike RBD and performed “forward folding” with DeepAb to assess their potential to fold into the designed structures. Ultimately, our results identified one designed Ab variable region, P1.D1, as a particularly promising candidate for experimental testing. This effort puts forth a computational workflow for the de novo design and evaluation of Abs that can quickly be adapted to target spike epitopes of emerging SARS-CoV-2 variants or other antigenic targets.

59 BASIC BIOLOGICAL SCIENCES↗

VTAnDeM: A python toolkit for simultaneously visualizing phase stability, defect energetics, and carrier concentrations of materials

Phase stability, defect formation energies, and carrier concentrations are closely interrelated features of semiconductors. Due to their joint dependence on the multidimensional chemical potential space, it is challenging to quantitatively establish patterns between these quantities in a given semiconductor, especially when the semiconductor is comprised of multiple elements. To enable synchronous visualization and analysis of these complementary material properties and their interdependence, we developed the Visualization Toolkit for Analyzing Defects in Materials (VTAnDeM). This python-based toolkit allows users to interactively explore how defect formation energies and carrier concentrations vary across the composition and chemical potential spaces of multicomponent semiconductors. Here, we illustrate the computational workflow that employs VTAnDeM as a post-processing tool for first-principles calculations and describe the data organization and theory underlying the visualization scheme. Furthermore, we believe that this software will serve as a useful tool for simultaneously visualizing the often complex and non-intuitive chemical potential – defect – carrier concentration phase space of semiconductors.

36 MATERIALS SCIENCE↗

Comparative Analysis via CFD Simulation on the Impact of Graphite Anode Morphologies on the Discharge of a Lithium-Ion Battery

The morphology of electrode materials plays a crucial role in determining the performance of lithium-ion batteries. Traditional computational models often simplify graphite flakes as uniformly sized spheres, which limits their predictive accuracy. In this study, we present a computational workflow that overcomes these limitations by incorporating a more realistic representation of graphite morphologies. This workflow is designed to be flexible and reproducible, enabling efficient evaluation of electrochemical performance across diverse material structures. By exploring different graphite morphologies, our approach accelerates the optimization of material preparation techniques and processing conditions. Our findings reveal that incorporating greater morphological complexity leads to significant deviations from classical model predictions. Instead, our refined model offers a more accurate representation of battery discharge behavior, closely aligning with experimental data. This improvement underscores the importance of detailed morphological descriptions in advancing battery design and performance assessments. To promote accessibility and reproducibility, we provide the developed code for seamless integration with the COMSOL API, allowing researchers to implement and adapt it easily. This computational framework serves as a valuable tool for investigating the impact of graphite morphology on battery performance, bridging the gap between theoretical modeling and experimental validation to enhance lithium-ion battery technology.

25 ENERGY STORAGE↗

Database of ab initio L-edge X-ray absorption near edge structure

Abstract The L-edge X-ray Absorption Near Edge Structure (XANES) is widely used in the characterization of transition metal compounds. Here, we report the development of a database of computed L-edge XANES using the multiple scattering theory-based FEFF9 code. The initial release of the database contains more than 140,000 L-edge spectra for more than 22,000 structures generated using a high-throughput computational workflow. The data is disseminated through the Materials Project and addresses a critical need for L-edge XANES spectra among the research community.

42 ENGINEERING↗

A structural homology approach to identify potential cross-reactive antibody responses following SARS-CoV-2 infection

Abstract The emergence of the novel SARS-CoV-2 virus is the most important public-health issue of our time. Understanding the diverse clinical presentations of the ensuing disease, COVID-19, remains a critical unmet need. Here we present a comprehensive listing of the diverse clinical indications associated with COVID-19. We explore the theory that anti-SARS-CoV-2 antibodies could cross-react with endogenous human proteins driving some of the pathologies associated with COVID-19. We describe a novel computational approach to estimate structural homology between SARS-CoV-2 proteins and human proteins. Antibodies are more likely to interrogate 3D-structural epitopes than continuous linear epitopes. This computational workflow identified 346 human proteins containing a domain with high structural homology to a SARS-CoV-2 Wuhan strain protein. Of these, 102 proteins exhibit functions that could contribute to COVID-19 clinical pathologies. We present a testable hypothesis to delineate unexplained clinical observations vis-à-vis COVID-19 and a tool to evaluate the safety-risk profile of potential COVID-19 therapies.

60 APPLIED LIFE SCIENCES↗

Modeling antiphase boundary energies of Ni 3 Al-based alloys using automated density functional theory and machine learning

Antiphase boundaries (APBs) are planar defects that play a critical role in strengthening Ni-based superalloys, and their sensitivity to alloy composition offers a flexible tuning parameter for alloy design. Here, we report a computational workflow to enable the development of sufficient data to train machine-learning (ML) models to automate the study of the effect of composition on the (111) APB energy in Ni 3 Al-based alloys. We employ ML to leverage this wealth of data and identify several physical properties that are used to build predictive models for the APB energy that achieve a cross-validation error of 0.033 J m –2 . We demonstrate the transferability of these models by predicting APB energies in commercial superalloys. Moreover, our use of physically motivated features such as the ordering energy and stoichiometry-based features opens the way to using existing materials properties databases to guide superalloy design strategies to maximize the APB energy.

36 MATERIALS SCIENCE↗

Symmetry is the Key to the Design of Reticular Frameworks

De novo prediction of reticular framework structures is a challenging task for chemists and materials scientists. Herein, a computational workflow that predicts a list of possible reticular frameworks based on only the connectivity and symmetry of node and linker building blocks is presented. This list is ranked based on the occurrence of topologies in known structures, thus providing a manageable number of structures that can be optimized using density functional theory, and inform future experiments. This workflow is broadly applicable, correctly predicts known reticular materials, and furthermore identifies novel unknown phases for some systems.

COF↗

Unveiling User Behavior on Summit Login Nodes as a User

We observe and analyze usage of the login nodes of the leadership class Summit supercomputer from the perspective of an ordinary user—not a system administrator—by periodically sampling user activities (job queues, running processes, etc.) for two full years (2020–2021). Our findings unveil key usage patterns that evidence misuse of the system, including gaming the policies, impairing I/O performance, and using login nodes as a sole computing resource. Our analysis highlights observed patterns for the execution of complex computations (workflows), which are key for processing large-scale applications.

Wilkinson, Sean↗

A2SD: Accelerating Scientific Innovation Through Autonomous Discovery Systems

The 2025 Advancing Autonomous Scientific Discovery (A2SD) workshop convened researchers from academia, national laboratories, and industry to explore the transformative role of autonomy in scientific discovery. The workshop highlighted a convergence of artificial intelligence, robotics, and computational workflows into autonomous systems capable of accelerating the scientific process. Presentations and discussions spanned autonomous experimentation, intelligent workflow orchestration, digital twins, and agent-based systems for managing complex research ecosystems. Key challenges discussed included interoperability across heterogeneous infrastructures, near real-time data management under FAIR principles, reproducibility, and the integration of human oversight. The workshop also emphasized the need for modular software interfaces, federated learning models, and education initiatives to support a next-generation scientific workforce.

Taufer, Michela [University of Tennessee, Knoxvill↗

Discovery, characterization, and application of chromosomal integration sites in the hyperthermophilic archaeon Sulfolobus islandicus

Sulfolobus islandicus , an emerging archaeal model organism, offers unique advantages for metabolic engineering and synthetic biology applications owing to its ability to thrive in extreme environments. Although several genetic tools have been established for this organism, the lack of well-characterized chromosomal integration sites has limited its potential as a cellular factory. Here, in this work, we systematically identified and characterized 13 artificial CRISPR RNAs targeting eight integration sites in S. islandicus using the CRISPR-COPIES pipeline and a multi-omics-informed computational workflow. We leveraged the endogenous CRISPR-Cas system to integrate the reporter gene lacS and validated heterologous expression through a β-galactosidase assay, revealing significant positional effects. As a proof of concept, we utilized these sites to genetically manipulate lipid ether composition by overexpressing glycerol dibiphytanyl glycerol tetraether (GDGT) ring synthase B (GrsB). This study expands the genetic toolbox for S. islandicus and advances its potential as a robust platform for archaeal synthetic biology and industrial biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence Modulates Polypeptoid Hydration Water Structure and Dynamics

We use molecular dynamics simulations to investigate the effect of polypeptoid sequence on the structure and dynamics of its hydration waters. Polypeptoids provide an excellent platform to study small-molecule hydration in disordered polymers, as they can be precisely synthesized with a variety of sidechain chemistries. We examine water behavior near a set of peptoid oligomers in which the number and placement of nonpolar versus polar sidechains are systematically varied. To do this, we leverage a new computational workflow enabling accurate sampling of polypeptoid conformations. We find that the hydration waters are less dense, are more tetrahedral, and have slower dynamics compared to bulk water. The magnitude of these shifts increases with the number of nonpolar groups. Here, we also find that shifts in the water structure and dynamics are strongly correlated, suggesting that experimental insight into the dynamics of hydration water obtained by Overhauser dynamic nuclear polarization (ODNP) also contains information about water structural properties. We then demonstrate the ability of ODNP to probe site-specific dynamics of hydration water near these model peptoid systems.

36 MATERIALS SCIENCE↗

Off-Equilibrium Reactivity of Boron-Enriched Metal Diboride Surfaces in Electroreduction Conditions

Boron-based materials, featuring B-dependent reactivity and diverse phases, are emerging as promising catalyst systems. However, the catalytic mechanism on many borides remains poorly understood due to complex surface reconstructions under reaction conditions. Here, we investigate the MoB 2 surface in conditions of hydrogen evolution reaction in acidic media, using grand canonical global optimization, grand canonical density functional theory, ab initio molecular dynamics, free energy surface sampling, and an analytical model for electrochemical barrier evaluation. We propose a boron-enrichment strategy to tune the surface reactivity of the hexagonal face of MoB 2 . We reveal the dynamic nature of the B-enriched surface under H coverage and kinetic trapping of the system in the metastable regime with an extensive examination of the deactivation pathways. The metastable center B site on B-enriched surfaces, featuring buckled-up configuration and a usual relaxation effect, is found to be highly active toward HER via the Volmer–Heyrovsky mechanism. In conclusion, this work demonstrates how off-equilibrium behaviors can arise from the interplay between adsorbate coverage and surface reconstruction on a seemingly simple surface, and we present a theoretical framework and computational workflows to address these behaviors, along with other realistic complexities, in kinetics simulations.

Adsorption↗

PeakDecoder enables machine learning-based metabolite annotation and accurate profiling in multidimensional mass spectrometry measurements

Multidimensional measurements using state-of-the-art separations and mass spectrometry provide advantages in untargeted metabolomics analyses for studying biological and environmental bio-chemical processes. However, the lack of rapid analytical methods and robust algorithms for these heterogeneous data has limited its application. Here, we develop and evaluate a sensitive and high-throughput analytical and computational workflow to enable accurate metabolite profiling. Our workflow combines liquid chromatography, ion mobility spectrometry and data-independent acquisition mass spectrometry with PeakDecoder, a machine learning-based algorithm that learns to distinguish true co-elution and co-mobility from raw data and calculates metabolite identification error rates. We apply PeakDecoder for metabolite profiling of various engineered strains of Aspergillus pseudoterreus, Aspergillus niger, Pseudomonas putida and Rhodosporidium toruloides. Results, validated manually and against selected reaction monitoring and gas-chromatography platforms, show that 2683 features could be confidently annotated and quantified across 116 microbial sample runs using a library built from 64 standards.

59 BASIC BIOLOGICAL SCIENCES↗

Strategies to search for two-dimensional materials with long spin qubit coherence time

Two-dimensional (2D) materials that can host qubits with long spin coherence time (T 2 ) have the distinct advantage of integrating easily with existing microelectronic and photonic platforms, making them attractive for designing novel quantum devices with enhanced performance. However, the relative lack of 2D materials as spin qubit hosts, as well as appropriate substrates that can help maintain long T 2 , necessitates a strategy to search for candidates with robust spin coherence. Here, we develop a high-throughput computational workflow to predict the nuclear spin bath-driven qubit decoherence and T 2 in 2D materials and heterostructures. We initially screen 1172 2D materials and find 189 monolayers with T 2 > 1 ms, higher than that of naturally-abundant diamond. We then construct 1554 lattice-commensurate heterostructures between high-T 2 2D materials and select 3D substrates, and we find that T 2 is generally lower in a heterostructure than in the bare 2D host material; however, low-noise substrates (such as CeO 2 and CaO) can help maintain high T 2 . To further accelerate the material screening effort, we derive analytical models that enable rapid predictions of T 2 for 2D materials and heterostructures. The models offer a simple, yet quantitative, way to determine the relative contributions to decoherence from the nuclear spin baths of the 2D host and substrate in a heterostructural system. By developing a high-throughput workflow and analytical models, we expand the genome of 2D materials and their spin coherence times for the development of spin qubit platforms.

Toriyama, Michael Y. [Argonne National Laboratory ↗

ROOT’s RNTuple I/O Subsystem: The Path to Production

The RNTuple I/O subsystem is ROOT’s future event data file format and access API. It is driven by the expected data volume increase at upcoming HEP experiments, e.g. at the HL-LHC, and recent opportunities in the storage hardware and software landscape such as NVMe drives and distributed object stores. RNTuple is a redesign of the TTree binary format and API and has shown to deliver substantially faster data throughput and better data compression both compared to TTree and to industry standard formats. In order to let HENP computing workflows benefit from RNTuple’s superior performance, however, the I/O stack needs to connect efficiently to the rest of the ecosystem, from grid storage to (distributed) analysis frameworks to (multithreaded) experiment frameworks for reconstruction and ntuple derivation. With the RNTuple binary format soon arriving at its first production release, we present RNTuple’s feature set, integration efforts, and its performance impact on the time-to-solution. We show the latest performance figures of RDataFrame analysis code of realistic complexity, comparing RNTuple and TTree as data sources. We discuss RNTuple’s approach to functionality critical to the HENP I/O (such as multithreaded writes, fast data merging, schema evolution) and we provide an outlook on the road to its use in production.

Blomer, Jakob↗