Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific AI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Cognitive analysis of metabolomics data for systems biology

Cognitive computing is revolutionizing the way big data are processed and integrated, with artificial intelligence (AI) natural language processing (NLP) platforms helping researchers to efficiently search and digest the vast scientific literature. Most available platforms have been developed for biomedical researchers, but new NLP tools are emerging for biologists in other fields and an important example is metabolomics. NLP provides literature-based contextualization of metabolic features that decreases the time and expert-level subject knowledge required during the prioritization, identification and interpretation steps in the metabolomics data analysis pipeline. Here, we describe and demonstrate four workflows that combine metabolomics data with NLP-based literature searches of scientific databases to aid in the analysis of metabolomics data and their biological interpretation. Additionally, the four procedures can be used in isolation or consecutively, depending on the research questions. The first, used for initial metabolite annotation and prioritization, creates a list of metabolites that would be interesting for follow-up. The second workflow finds literature evidence of the activity of metabolites and metabolic pathways in governing the biological condition on a systems biology level. The third is used to identify candidate biomarkers, and the fourth looks for metabolic conditions or drug-repurposing targets that the two diseases have in common. The protocol can take 1–4 h or more to complete, depending on the processing time of the various software used.

59 BASIC BIOLOGICAL SCIENCES↗

Scientific Data Management Beyond Traditional Computing Boundaries

Scientific data management is undergoing a fundamental transformation driven by the convergence of artificial intelligence (AI)/machine learning workflows, distributed computing and storage environments, and exponential data growth. Here, we analyze how these developments address current limitations while enabling new capabilities for cross-facility collaboration and AI-driven research.

Widener, Patrick [Oak Ridge National Laboratory (O↗

Towards interpretable Cryo-EM: disentangling latent spaces of molecular conformations

Molecules are essential building blocks of life and their different conformations (i.e., shapes) crucially determine the functional role that they play in living organisms. Cryogenic Electron Microscopy (cryo-EM) allows for acquisition of large image datasets of individual molecules. Recent advances in computational cryo-EM have made it possible to learn latent variable models of conformation landscapes. However, interpreting these latent spaces remains a challenge as their individual dimensions are often arbitrary. The key message of our work is that this interpretation challenge can be viewed as an Independent Component Analysis (ICA) problem where we seek models that have the property of identifiability. That means, they have an essentially unique solution, representing a conformational latent space that separates the different degrees of freedom a molecule is equipped with in nature. Thus, we aim to advance the computational field of cryo-EM beyond visualizations as we connect it with the theoretical framework of (nonlinear) ICA and discuss the need for identifiable models, improved metrics, and benchmarks. Moving forward, we propose future directions for enhancing the disentanglement of latent spaces in cryo-EM, refining evaluation metrics and exploring techniques that leverage physics-based decoders of biomolecular systems. Moreover, we discuss how future technological developments in time-resolved single particle imaging may enable the application of nonlinear ICA models that can discover the true conformation changes of molecules in nature. The pursuit of interpretable conformational latent spaces will empower researchers to unravel complex biological processes and facilitate targeted interventions. This has significant implications for drug discovery and structural biology more broadly. More generally, latent variable models are deployed widely across many scientific disciplines. Thus, the argument we present in this work has much broader applications in AI for science if we want to move from impressive nonlinear neural network models to mathematically grounded methods that can help us learn something new about nature.

59 BASIC BIOLOGICAL SCIENCES↗

Automation of planetary spacecraft

The development of autonomous spacecraft from 1960 to the present is traced within a framework of the definitions and measures of the level of autonomy. The attainment of milestones in the level of autonomy in spacecraft guidance and control is described in terms of the Mariner, Viking, Voyager, Galileo and Mark II (under development) spacecraft. The constant interplay between the definition of scientific mission goals and available technological capabilities is explored, along with current efforts to implement AI techniques and advanced software in spacecraft to allow reliable functioning in stressful conditions.

Varsi, G.↗

Second Conference on Artificial Intelligence for Space Applications

The proceedings of the conference are presented. This second conference on Artificial Intelligence for Space Applications brings together a diversity of scientific and engineering work and is intended to provide an opportunity for those who employ AI methods in space applications to identify common goals and to discuss issues of general interest in the AI community.

Dollman, Thomas↗

Learning from learning machines: improving the predictive power of energy-water-land nexus models with insights from complex measured and simulated data

Focal Area(s): Insights gleaned from complex data (observed and simulated) using AI; Predictive modeling through the use of AI techniques and AI-derived model components, including physics- and knowledge-informed models; Energy-water-land nexus and integrated energy systems – models of MultiSector dynamics. Science Challenge: Scientific communities are in need of tools for the computational integration of physics-based models, experimental data, and empirical/observational studies across a broad range of temporal and spatial scales to explore and meet EESSD grand challenges. We require AI algorithms for the discovery of process-drivers in the Earth-energy-human system and to link them with complex measured and simulated data. We aim to understand the energy-water-land nexus, particularly under extreme forcing scenarios and rare events, leveraging scale-aware AI process models, including probabilistic uncertainties, and benefiting from the emerging 5G-enabled landscape-scale sensing and edge computing capabilities. These models are needed to identify instabilities and tipping points that manifest extreme system behaviors with consequences for integrated energy systems and the environment. This work requires fundamental advances in uncertainty quantification, in particular, to identify and model unlikely but catastrophic outliers. Specifically, we need models that are interpretable to domain scientists and, essentially, explainable to public and private stake holders.

54 ENVIRONMENTAL SCIENCES↗

Data-Enabled Fusion Technology (Final Scientific/Technical Report)

Advancing Scientific Understanding in Fusion Energy and Machine Learning This research represented a significant step forward in machine learning (ML) applications for fusion energy experiments. The project integrated advanced data-driven modeling, optimization techniques, and artificial intelligence to enhance the predictive capabilities and operational efficiency of plasma-based fusion systems. Specifically, tasks focused on ML-enhanced diagnostics, operator guidance tools, and predictive modeling helped improve the ability to interpret complex fusion experiments. Key areas of advancement included: 1) data-driven plasma control, i.e., using ML algorithms to optimize experimental conditions and classify plasma behaviors based on historical data; 2) spectroscopy and diagnostics, i.e., applying AI models to extract previously inaccessible insights from experimental spectroscopy data; and 3) configuration mapping and operator guidance, i.e., developing a predictive framework to assist scientists in identifying the most effective experimental parameters, reducing reliance on manual adjustments. By refining these ML-driven techniques, the project contributed to the broader scientific community’s understanding of plasma dynamics and fusion energy viability. Technical Effectiveness and Economic Feasibility The methods investigated demonstrated high technical effectiveness, as reflected in milestones assessing the predictive accuracy, performance, and optimization of fusion configurations. The development of an Operator Guidance Tool (OGT), for example, led to more precise control of plasma conditions by learning from experimental data and offering real-time adjustments. From an economic standpoint, DeFT provided: 1) the ability to reduce trial-and-error experimentation, which lowered operational costs; 2) improved data interpretation methods, which enabled more efficient resource allocation in large-scale fusion research projects; and 3) the automation of key diagnostic tasks, which reduced manual labor and human error, increasing overall efficiency. 13 The final assessments of predictive models and optimization strategies demonstrated that these approaches were scalable and could be implemented across multiple fusion energy research programs. Public Benefit and Societal Impact This project contributed directly to the broader goal of achieving sustainable and commercially viable fusion energy, which had profound implications for clean energy production and climate change mitigation. The integration of AI-driven solutions into fusion research: 1) sped up scientific discovery, accelerating progress towards achieving energy breakthroughs; 2) reduced the cost of experimentation, making fusion research more accessible; and 3) provided a framework for future AI applications in high-energy physics, benefiting adjacent fields like space exploration, material science, and renewable energy. Additionally, by fostering collaborations between AI researchers and plasma physicists, this project promoted interdisciplinary innovation that could lead to broader applications beyond fusion research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Brochure on the 2024 ASCR Workshop on Energy-Efficient Computing for Science

Large-scale computing has enabled numerous scientific discoveries, including ground-breaking achievements facilitated by the US Department of Energy (DOE) supercomputers and advances in applied mathematics and computer science. While important advances were made in energy efficiency to enable exascale computing, continued efforts are needed to dramatically improve the energy efficiency of the next generation of high-performance computing (HPC) systems and, more broadly, AI data centers. Without substantial improvements in energy efficiency, the energy consumption associated with computing could become a limiting factor for future scientific discovery, national security, and technological advancement.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Modern Scientific Data Governance Framework

Science has entered the era of Big Data with new challenges related to data governance, stewardship, and management. The existing data governance practices must catch up to ensure proper data management. Existing data governance policies and stewardship best practices tend to be disconnected from operational data management practices and enforcement and mainly exist in well-meaning documents or reports. These governance policies are, at best, partially implemented and rarely monitored or audited. In addition, existing governance policies keep adding additional data management steps that require a human, ‘a data steward’, in the loop, and the cost of data management can no longer scale proportionately with the current and future increased data volume and complexity. The goal for developing an updated data governance framework is to modernize scientific data governance to the reality of Big data and align it with the current technology trends such as cloud computing and AI. The goals of this framework are two folds. One is to ensure thoroughness that the governance adequately covers the entire data life cycle. Two, provide a practical approach that offers a consistent and repeatable process for different projects. Three core principles ground this framework. First, focus on just enough governance and prevent data governance from becoming a roadblock toward the scientific process. Remove any unnecessary processes and steps. Second, automate data management steps where possible. Actively remove steps that require ‘human in the loop’ within the management process to be efficient and scale with increasing data. Third, all the processes should continually be optimized using quantified metrics to streamline the monitoring and auditing workflows.

Rahul Ramachandran↗

5G Enabled Energy Innovation: Advanced Wireless Networks for Science (Workshop Report)

Rapidly expanding, new telecommunications infrastructure based on 5G technologies will disrupt and transform how we design, build, operate, and optimize scientific infrastructure and the experiments and services enabled by that infrastructure, from continental-scale sensor networks to centralized scientific user facilities, from intelligent Internet of Things devices to supercomputers. Concurrently, 5G will introduce, or exacerbate, challenges related to protecting infrastructure and associated scientific data as well as to fully leveraging opportunities related to expanded infrastructure scale and complexity. The U.S. Department of Energy (DOE) Office of Science operates scientific infrastructure, supporting some of the nation’s most advanced intellectual discoveries, spanning the country and including 30 world-class user facilities from supercomputers to accelerators. Along with field experiments and remote observatories, every aspect of DOE’s scientific enterprise will be affected by 5G, which amounts to a complete renovation of the underpinnings of the nation’s information infrastructure. In this report we explore the scientific opportunities and new research challenges associated with 5G, ranging from scalability to heterogeneity to cybersecurity. The rapid commercial deployment of 5G opens the opportunity to rethink and reinvent DOE’s scientific infrastructure and experimentation, from intelligent sensor networks at unprecedented scales to a digital continuum of cyberinfrastructure spanning low-power sensors, high-performance computing embedded within and at the edge of the network, and DOE’s large-scale user instrument and computing facilities. New programming paradigms, workflow and data frameworks, and AI-based system design, operation, and autonomous adaptation and optimization will be necessary in order to exploit these new opportunities. Field deployments and centralized scientific instruments can also be revolutionized, moving (without traditional performance penalties) from wired to wireless connectivity for data and control systems, improving flexibility, and opening new sensing modalities, including the use of the 5G electromagnetic spectrum itself as an environmental probe. For DOE science, in contrast to commercial 5G applications and settings, devices will be deployed in extreme environments such as cryogenically cooled instrument control systems and in remote settings with harsh conditions, requiring the design of new materials for RF communication and edge processing to operate in these regimes. Concurrently, 5G infrastructure comprises both hardware and sophisticated software systems - currently closed and proprietary. The cybersecurity challenges to 5G-empowered reinvention mirror the complexity and variety of new 5G features, from virtualization to private network slices to ubiquitous access. Research is also needed in order to accelerate the development of secure and open 5G software infrastructure, reducing reliance on hardware and software produced outside the United States and providing the transparency and rigorous evaluation and testing afforded through open software. Twelve broad research thrusts are laid out in four chapters, with a companion fifth chapter (and three additional research thrusts) underscoring the needs and opportunities for an aggressive testbed program co-designed by networking experts and scientists involved in the 15 research thrusts. The urgency of undertaking this research is fueled by a global, accelerating deployment of new telecommunications infrastructure that is designed for entertainment and commercial applications - barely scratching the surface of what 5G can do to extend U.S. leadership in scientific discovery.

42 ENGINEERING↗

Machine Learning and Data Science to Advance Laboratory Earthquake Prediction and Illuminate the Mechanics of Precursors to Failure

Earthquakes represent one of our greatest natural hazards and in recent years human induced seismicity is adding to the threat. Even a modest improvement in the ability to forecast devastating large earthquakes or smaller shallow events associated with fluid injection could save thousands of lives and billions of dollars. Current efforts to forecast earthquakes are limited by knowledge of earthquake physics and hampered by a lack of reliable lab or field observations. However, recent work has provided a critical opportunity for advancement. We have found: 1) clear and consistent precursors prior to earthquake-like failure in the laboratory and 2) that lab earthquakes can be predicted using machine learning (ML). These works show that stick-slip failure events –the lab equivalent of earthquakes– are preceded by a cascade of micro-failure events that radiate elastic energy in a manner that foretells catastrophic failure. Remarkably, ML predicts the fault zone stress state, the failure time and in some cases the magnitude of lab earthquakes. In addition, the observations include clear precursors to failure in the form of changes in fault zone properties prior to lab earthquakes. Precursors have been observed in previous laboratory studies but their origin is poorly understood and their possible connection to ML based earthquake prediction is unknown. The work conducted under our project has dramatically expanded these efforts. We have developed an integrated data science approach to illuminate the physics of earthquake precursors and lab earthquake prediction. Our work has accelerated the development of ML, artificial intelligence (AI), and related data science approaches by providing massive data sets that are tightly connected to critical scientific problems and by bringing together leading subject matter experts and data scientists. Earthquake physics involves phenomena that are far from equilibrium. Our work has leveraged data science methods to illuminate these phenomena and investigate how they relate to earthquake prediction. In addition to a large database with many types of labeled events that is available to everyone, our work has advanced the fundamental understanding of seismic forecasting, earthquake physics, and fault rheology

58 GEOSCIENCES↗

Developing and Distributing HEP Software Stacks with Spack

The Computational Science and AI Directorate at Fermilab is using Spack to support the development efforts of a large number of scientific programmers, in many independent projects and experiments. While independent, these projects share many dependencies. They are typically under continuous and fairly rapid development. They have to support deployment on diverse hardware. This is a different context than is typical for the management of HPC software, where Spack was born. To support our community, we have created a model that enables users to develop code with greater efficiency than is possible with Spack’s current development facilities. In this talk we will present: - a brief introduction to the science we support (particle physics) - how the code we work with is naturally organized into several layers of packages - how we are using Spack to manage those layers - how we leverage the layering to provide efficient support for developers, using our Spack extension “MPD”. - some suggestions for changes or additions to Spack to make such work easier.

Knoepfel, Kyle J. [Fermilab]↗

Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics

The primary objective of this project is to strengthen the trustworthiness of AI systems by designing algorithms that make their internal decision-making processes more understandable to human users. This involves creating clear, interpretable explanations for AI decisions and developing metrics to assess these explanations' validity and reliability. Significant progress has been achieved through (i) developing symbolic explanations, (ii) generating meaningful interpretive insights, (iii) establishing accuracy and confidence metrics, and (iv) devising methods to evaluate the knowledge boundaries of AI models. To date, the research findings have been shared in peer-reviewed publications, with accompanying scientific and technical information (STI) detailed below.

97 MATHEMATICS AND COMPUTING↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

PowerModel-AI: A First On-the-Fly Machine-Learning Predictor for AC Power Flow Solutions

The real-time creation of machine-learning models via active or on-the-fly learning has attracted considerable interest across various scientific and engineering disciplines. These algorithms enable machines to build models autonomously while remaining operational. Through a series of query strategies, the machine can evaluate whether newly encountered data fall outside the scope of the existing training set. In this study, we introduce PowerModel-AI, an end-to-end machine learning software designed to accurately predict AC power flow solutions. We present detailed justifications for our model design choices and demonstrate that selecting the right input features effectively captures load flow decoupling inherent in power flow equations. Our approach incorporates on-the-fly learning, where power flow calculations are initiated only when the machine detects a need to improve the dataset in regions where the model’s suboptimal performance is based on specific criteria. Otherwise, the existing model is used for power flow predictions. This study includes analyses of five Texas A&M synthetic power grid cases, encompassing the 14-, 30-, 37-, 200-, and 500-bus systems. The training and test datasets were generated using PowerModels.jl, an open-source power flow solver/optimizer developed at Los Alamos National Laboratory, NM, USA.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Agnostic capture of pathogens for the detection and diagnostics of emerging threats

The continued emergence of pathogens, whether novel, re-emerging, or engineered, poses a persistent global biosecurity and public health challenge. Recent outbreaks, including COVID-19, Lassa fever, Marburg virus, mpox, and avian influenza, underscore the urgent need for robust systems that enable rapid surveillance, early diagnosis, and timely countermeasures before widespread human transmission occurs. In this article, we focus on early detection technologies and systematically evaluate current diagnostic and sensing modalities. We highlight sequencing and spectroscopy as two complementary approaches capable of providing broad, agnostic detection and rich biological insight. Our analysis emphasizes that scientific innovation alone is insufficient: effective preparedness also requires improved data curation, integration, and sharing to build AI-ready resources that accelerate future responses. We argue for coordinated advances in both technological capabilities and supporting infrastructure to enable the rapid identification and characterization of emerging pathogens and to fully leverage modern science against evolving infectious threats.

Environmental health↗

ToPolyAgent: AI agents for coarse-grained bead-spring topological polymer simulations

We introduce ToPolyAgent, a multi-agent AI framework for performing coarse-grained molecular dynamics (MD) simulations of topological polymers through natural language instructions. By integrating large language models (LLMs) with domain-specific computational tools, ToPolyAgent supports both interactive and autonomous simulation workflows across diverse polymer architectures, including linear, ring, brush, and star polymers, as well as dendrimers. The system consists of four LLM-powered agents: a Config Agent for generating initial polymer–solvent configurations, a Simulation Agent for executing LAMMPS-based MD simulations and conformational analyses, a Report Agent for compiling markdown reports, and a Workflow Agent for streamlined autonomous operations. Interactive mode incorporates user feedback loops for iterative refinements, while autonomous mode enables end-to-end task execution from detailed prompts. We demonstrate ToPolyAgent's versatility through case studies involving diverse polymer architectures under varying solvent conditions, thermostats, and simulation lengths. Furthermore, we highlight its potential as a research assistant by directing it to investigate the effect of interaction parameters on the linear polymer conformation, and the influence of grafting density on the persistence length of the brush polymer. By coupling natural language interfaces with rigorous simulation tools, ToPolyAgent lowers barriers to complex computational workflows and advances AI-driven materials discovery in polymer science. It lays the foundation for autonomous and extensible multi-agent scientific research ecosystems.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

Edge at the Pier: EPCAPE Software-Defined Sensing Field Campaign Report

The Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) was aimed to enhance the understanding of cloud and aerosol properties in the region surrounding La Jolla, California. To address challenges in data collection and processing from various instruments, an edge computing device known as Waggle Sage Node (WSN) was deployed at the Ellen Browning Scripps Memorial Pier. WSN is a distributed-sensing platform designed to collect and analyze environmental data at the edge. Sage is a multi-agency-supported project that designs and builds a new kind of national-scale reusable cyberinfrastructure to enable artificial intelligence (AI) at the edge based on the Waggle platform. Sponsors include the U.S. Department of Energy (DOE) Advanced Scientific Computing Research (ASCR), DOE National Nuclear Security Administration (NNSA), DOE Biological and Environmental Research (BER) through DOE Artificial Intelligence for Earth System Predictability (AI4ESP), Argonne Laboratory-Directed Research and Development (LDRD). Sage (https://sagecontinuum.org/) is funded as a National Science Foundation Mid-Scale Research Infrastructure (MSRI) project (https://www.nsf.gov/awardsearch/showAward?AWD_ID=1935984). This robust, multi-architecture edge computing platform facilitated environmental monitoring during the campaign. This report details the scientific objectives, deployment process, and key results of integrating Waggle into the EPCAPE field campaign.

54 ENVIRONMENTAL SCIENCES↗