Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compute workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Turbo‐charging crop improvement: harnessing multiplex editing for polygenic trait engineering and beyond

Multiplex CRISPR editing has emerged as a transformative platform for plant genome engineering, enabling the simultaneous targeting of multiple genes, regulatory elements, or chromosomal regions. This approach is effective for dissecting gene family functions, addressing genetic redundancy, engineering polygenic traits, and accelerating trait stacking and de novo domestication. Its applications now extend beyond standard gene knockouts to include epigenetic and transcriptional regulation, chromosomal engineering, and transgene‐free editing. These capabilities are advancing crop improvement not only in annual species but also in more complex systems such as polyploids, undomesticated wild relatives, and species with long generation times. At the same time, multiplex editing presents technical challenges, including complex construct design and the need for robust, scalable mutation detection. We discuss current toolkits and recent innovations in vector architecture, such as promoter and scaffold engineering, that streamline workflows and enhance editing efficiency. High‐throughput sequencing technologies, including long‐read platforms, are improving the resolution of complex editing outcomes such as structural rearrangements—often missed by standard genotyping—when targeting repetitive or tandemly spaced loci. To fully realize the potential of multiplex genome engineering, there is growing demand for user‐friendly, synthetic biology‐compatible, and scalable computational workflows for gRNA design, construct assembly, and mutation analysis. Experimentally validated inducible or tissue‐specific promoters are also highly desirable for achieving spatiotemporal control. As these tools continue to evolve, multiplex CRISPR editing is poised to become a foundational technology of next‐generation crop improvement to address challenges in agriculture, sustainability, and climate resilience.

59 BASIC BIOLOGICAL SCIENCES↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

AI-Accelerated Design of Targeted Covalent Inhibitors for SARS-CoV-2

Direct-acting antivirals for the treatment of the COVID-19 pandemic caused by the SARS-CoV-2 virus are needed to complement vaccination efforts. Given the ongoing emergence of new variants, automated experimentation, and active learning based fast workflows for antiviral lead discovery remain critical to our ability to address the pandemic’s evolution in a timely manner. While several such pipelines have been introduced to discover candidates with noncovalent interactions with the main protease (M pro ), here we developed a closed-loop artificial intelligence pipeline to design electrophilic warhead-based covalent candidates. Here, this work introduces a deep learning-assisted automated computational workflow to introduce linkers and an electrophilic “warhead” to design covalent candidates and incorporates cutting-edge experimental techniques for validation. Using this process, promising candidates in the library were screened, and several potential hits were identified and tested experimentally using native mass spectrometry and fluorescence resonance energy transfer (FRET)-based screening assays. We identified four chloroacetamide-based covalent inhibitors of M pro with micromolar affinities (K I of 5.27 μM) using our pipeline. Experimentally resolved binding modes for each compound were determined using room-temperature X-ray crystallography, which is consistent with the predicted poses. The induced conformational changes based on molecular dynamics simulations further suggest that the dynamics may be an important factor to further improve selectivity, thereby effectively lowering KI and reducing toxicity. These results demonstrate the utility of our modular and data-driven approach for potent and selective covalent inhibitor discovery and provide a platform to apply it to other emerging targets.

60 APPLIED LIFE SCIENCES↗

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

De novo design and Rosetta-based assessment of high-affinity antibody variable regions (Fv) against the SARS-CoV -2 spike receptor binding domain ( RBD )

The continued emergence of new SARS-CoV-2 variants has accentuated the growing need for fast and reliable methods for the design of potentially neutralizing antibodies (Abs) to counter immune evasion by the virus. Here, we report on the de novo computational design of high-affinity Ab variable regions (Fv) through the recombination of VDJ genes targeting the most solvent-exposed hACE2-binding residues of the SARS-CoV-2 spike receptor binding domain (RBD) protein using the software tool OptMAVEn-2.0. Subsequently, we carried out computational affinity maturation of the designed variable regions through amino acid substitutions for improved binding with the target epitope. Immunogenicity of designs was restricted by preferring designs that match sequences from a 9-mer library of “human Abs” based on a human string content score. We generated 106 different antibody designs and reported in detail on the top five that trade-off the greatest computational binding affinity for the RBD with human string content scores. We further describe computational evaluation of the top five designs produced by OptMAVEn-2.0 using a Rosetta-based approach. We used Rosetta SnugDock for local docking of the designs to evaluate their potential to bind the spike RBD and performed “forward folding” with DeepAb to assess their potential to fold into the designed structures. Ultimately, our results identified one designed Ab variable region, P1.D1, as a particularly promising candidate for experimental testing. This effort puts forth a computational workflow for the de novo design and evaluation of Abs that can quickly be adapted to target spike epitopes of emerging SARS-CoV-2 variants or other antigenic targets.

59 BASIC BIOLOGICAL SCIENCES↗

VTAnDeM: A python toolkit for simultaneously visualizing phase stability, defect energetics, and carrier concentrations of materials

Phase stability, defect formation energies, and carrier concentrations are closely interrelated features of semiconductors. Due to their joint dependence on the multidimensional chemical potential space, it is challenging to quantitatively establish patterns between these quantities in a given semiconductor, especially when the semiconductor is comprised of multiple elements. To enable synchronous visualization and analysis of these complementary material properties and their interdependence, we developed the Visualization Toolkit for Analyzing Defects in Materials (VTAnDeM). This python-based toolkit allows users to interactively explore how defect formation energies and carrier concentrations vary across the composition and chemical potential spaces of multicomponent semiconductors. Here, we illustrate the computational workflow that employs VTAnDeM as a post-processing tool for first-principles calculations and describe the data organization and theory underlying the visualization scheme. Furthermore, we believe that this software will serve as a useful tool for simultaneously visualizing the often complex and non-intuitive chemical potential – defect – carrier concentration phase space of semiconductors.

36 MATERIALS SCIENCE↗

Comparative Analysis via CFD Simulation on the Impact of Graphite Anode Morphologies on the Discharge of a Lithium-Ion Battery

The morphology of electrode materials plays a crucial role in determining the performance of lithium-ion batteries. Traditional computational models often simplify graphite flakes as uniformly sized spheres, which limits their predictive accuracy. In this study, we present a computational workflow that overcomes these limitations by incorporating a more realistic representation of graphite morphologies. This workflow is designed to be flexible and reproducible, enabling efficient evaluation of electrochemical performance across diverse material structures. By exploring different graphite morphologies, our approach accelerates the optimization of material preparation techniques and processing conditions. Our findings reveal that incorporating greater morphological complexity leads to significant deviations from classical model predictions. Instead, our refined model offers a more accurate representation of battery discharge behavior, closely aligning with experimental data. This improvement underscores the importance of detailed morphological descriptions in advancing battery design and performance assessments. To promote accessibility and reproducibility, we provide the developed code for seamless integration with the COMSOL API, allowing researchers to implement and adapt it easily. This computational framework serves as a valuable tool for investigating the impact of graphite morphology on battery performance, bridging the gap between theoretical modeling and experimental validation to enhance lithium-ion battery technology.

25 ENERGY STORAGE↗

A structural homology approach to identify potential cross-reactive antibody responses following SARS-CoV-2 infection

Abstract The emergence of the novel SARS-CoV-2 virus is the most important public-health issue of our time. Understanding the diverse clinical presentations of the ensuing disease, COVID-19, remains a critical unmet need. Here we present a comprehensive listing of the diverse clinical indications associated with COVID-19. We explore the theory that anti-SARS-CoV-2 antibodies could cross-react with endogenous human proteins driving some of the pathologies associated with COVID-19. We describe a novel computational approach to estimate structural homology between SARS-CoV-2 proteins and human proteins. Antibodies are more likely to interrogate 3D-structural epitopes than continuous linear epitopes. This computational workflow identified 346 human proteins containing a domain with high structural homology to a SARS-CoV-2 Wuhan strain protein. Of these, 102 proteins exhibit functions that could contribute to COVID-19 clinical pathologies. We present a testable hypothesis to delineate unexplained clinical observations vis-à-vis COVID-19 and a tool to evaluate the safety-risk profile of potential COVID-19 therapies.

60 APPLIED LIFE SCIENCES↗

Modeling antiphase boundary energies of Ni 3 Al-based alloys using automated density functional theory and machine learning

Antiphase boundaries (APBs) are planar defects that play a critical role in strengthening Ni-based superalloys, and their sensitivity to alloy composition offers a flexible tuning parameter for alloy design. Here, we report a computational workflow to enable the development of sufficient data to train machine-learning (ML) models to automate the study of the effect of composition on the (111) APB energy in Ni 3 Al-based alloys. We employ ML to leverage this wealth of data and identify several physical properties that are used to build predictive models for the APB energy that achieve a cross-validation error of 0.033 J m –2 . We demonstrate the transferability of these models by predicting APB energies in commercial superalloys. Moreover, our use of physically motivated features such as the ordering energy and stoichiometry-based features opens the way to using existing materials properties databases to guide superalloy design strategies to maximize the APB energy.

36 MATERIALS SCIENCE↗

Symmetry is the Key to the Design of Reticular Frameworks

De novo prediction of reticular framework structures is a challenging task for chemists and materials scientists. Herein, a computational workflow that predicts a list of possible reticular frameworks based on only the connectivity and symmetry of node and linker building blocks is presented. This list is ranked based on the occurrence of topologies in known structures, thus providing a manageable number of structures that can be optimized using density functional theory, and inform future experiments. This workflow is broadly applicable, correctly predicts known reticular materials, and furthermore identifies novel unknown phases for some systems.

COF↗

Unveiling User Behavior on Summit Login Nodes as a User

We observe and analyze usage of the login nodes of the leadership class Summit supercomputer from the perspective of an ordinary user—not a system administrator—by periodically sampling user activities (job queues, running processes, etc.) for two full years (2020–2021). Our findings unveil key usage patterns that evidence misuse of the system, including gaming the policies, impairing I/O performance, and using login nodes as a sole computing resource. Our analysis highlights observed patterns for the execution of complex computations (workflows), which are key for processing large-scale applications.

Wilkinson, Sean↗

A2SD: Accelerating Scientific Innovation Through Autonomous Discovery Systems

The 2025 Advancing Autonomous Scientific Discovery (A2SD) workshop convened researchers from academia, national laboratories, and industry to explore the transformative role of autonomy in scientific discovery. The workshop highlighted a convergence of artificial intelligence, robotics, and computational workflows into autonomous systems capable of accelerating the scientific process. Presentations and discussions spanned autonomous experimentation, intelligent workflow orchestration, digital twins, and agent-based systems for managing complex research ecosystems. Key challenges discussed included interoperability across heterogeneous infrastructures, near real-time data management under FAIR principles, reproducibility, and the integration of human oversight. The workshop also emphasized the need for modular software interfaces, federated learning models, and education initiatives to support a next-generation scientific workforce.

Taufer, Michela [University of Tennessee, Knoxvill↗

Discovery, characterization, and application of chromosomal integration sites in the hyperthermophilic archaeon Sulfolobus islandicus

Sulfolobus islandicus , an emerging archaeal model organism, offers unique advantages for metabolic engineering and synthetic biology applications owing to its ability to thrive in extreme environments. Although several genetic tools have been established for this organism, the lack of well-characterized chromosomal integration sites has limited its potential as a cellular factory. Here, in this work, we systematically identified and characterized 13 artificial CRISPR RNAs targeting eight integration sites in S. islandicus using the CRISPR-COPIES pipeline and a multi-omics-informed computational workflow. We leveraged the endogenous CRISPR-Cas system to integrate the reporter gene lacS and validated heterologous expression through a β-galactosidase assay, revealing significant positional effects. As a proof of concept, we utilized these sites to genetically manipulate lipid ether composition by overexpressing glycerol dibiphytanyl glycerol tetraether (GDGT) ring synthase B (GrsB). This study expands the genetic toolbox for S. islandicus and advances its potential as a robust platform for archaeal synthetic biology and industrial biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence Modulates Polypeptoid Hydration Water Structure and Dynamics

We use molecular dynamics simulations to investigate the effect of polypeptoid sequence on the structure and dynamics of its hydration waters. Polypeptoids provide an excellent platform to study small-molecule hydration in disordered polymers, as they can be precisely synthesized with a variety of sidechain chemistries. We examine water behavior near a set of peptoid oligomers in which the number and placement of nonpolar versus polar sidechains are systematically varied. To do this, we leverage a new computational workflow enabling accurate sampling of polypeptoid conformations. We find that the hydration waters are less dense, are more tetrahedral, and have slower dynamics compared to bulk water. The magnitude of these shifts increases with the number of nonpolar groups. Here, we also find that shifts in the water structure and dynamics are strongly correlated, suggesting that experimental insight into the dynamics of hydration water obtained by Overhauser dynamic nuclear polarization (ODNP) also contains information about water structural properties. We then demonstrate the ability of ODNP to probe site-specific dynamics of hydration water near these model peptoid systems.

36 MATERIALS SCIENCE↗

Off-Equilibrium Reactivity of Boron-Enriched Metal Diboride Surfaces in Electroreduction Conditions

Boron-based materials, featuring B-dependent reactivity and diverse phases, are emerging as promising catalyst systems. However, the catalytic mechanism on many borides remains poorly understood due to complex surface reconstructions under reaction conditions. Here, we investigate the MoB 2 surface in conditions of hydrogen evolution reaction in acidic media, using grand canonical global optimization, grand canonical density functional theory, ab initio molecular dynamics, free energy surface sampling, and an analytical model for electrochemical barrier evaluation. We propose a boron-enrichment strategy to tune the surface reactivity of the hexagonal face of MoB 2 . We reveal the dynamic nature of the B-enriched surface under H coverage and kinetic trapping of the system in the metastable regime with an extensive examination of the deactivation pathways. The metastable center B site on B-enriched surfaces, featuring buckled-up configuration and a usual relaxation effect, is found to be highly active toward HER via the Volmer–Heyrovsky mechanism. In conclusion, this work demonstrates how off-equilibrium behaviors can arise from the interplay between adsorbate coverage and surface reconstruction on a seemingly simple surface, and we present a theoretical framework and computational workflows to address these behaviors, along with other realistic complexities, in kinetics simulations.

Adsorption↗

PeakDecoder enables machine learning-based metabolite annotation and accurate profiling in multidimensional mass spectrometry measurements

Multidimensional measurements using state-of-the-art separations and mass spectrometry provide advantages in untargeted metabolomics analyses for studying biological and environmental bio-chemical processes. However, the lack of rapid analytical methods and robust algorithms for these heterogeneous data has limited its application. Here, we develop and evaluate a sensitive and high-throughput analytical and computational workflow to enable accurate metabolite profiling. Our workflow combines liquid chromatography, ion mobility spectrometry and data-independent acquisition mass spectrometry with PeakDecoder, a machine learning-based algorithm that learns to distinguish true co-elution and co-mobility from raw data and calculates metabolite identification error rates. We apply PeakDecoder for metabolite profiling of various engineered strains of Aspergillus pseudoterreus, Aspergillus niger, Pseudomonas putida and Rhodosporidium toruloides. Results, validated manually and against selected reaction monitoring and gas-chromatography platforms, show that 2683 features could be confidently annotated and quantified across 116 microbial sample runs using a library built from 64 standards.

59 BASIC BIOLOGICAL SCIENCES↗

Strategies to search for two-dimensional materials with long spin qubit coherence time

Two-dimensional (2D) materials that can host qubits with long spin coherence time (T 2 ) have the distinct advantage of integrating easily with existing microelectronic and photonic platforms, making them attractive for designing novel quantum devices with enhanced performance. However, the relative lack of 2D materials as spin qubit hosts, as well as appropriate substrates that can help maintain long T 2 , necessitates a strategy to search for candidates with robust spin coherence. Here, we develop a high-throughput computational workflow to predict the nuclear spin bath-driven qubit decoherence and T 2 in 2D materials and heterostructures. We initially screen 1172 2D materials and find 189 monolayers with T 2 > 1 ms, higher than that of naturally-abundant diamond. We then construct 1554 lattice-commensurate heterostructures between high-T 2 2D materials and select 3D substrates, and we find that T 2 is generally lower in a heterostructure than in the bare 2D host material; however, low-noise substrates (such as CeO 2 and CaO) can help maintain high T 2 . To further accelerate the material screening effort, we derive analytical models that enable rapid predictions of T 2 for 2D materials and heterostructures. The models offer a simple, yet quantitative, way to determine the relative contributions to decoherence from the nuclear spin baths of the 2D host and substrate in a heterostructural system. By developing a high-throughput workflow and analytical models, we expand the genome of 2D materials and their spin coherence times for the development of spin qubit platforms.

Toriyama, Michael Y. [Argonne National Laboratory ↗