Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Policy Considerations When Federating Facilities for Experimental and Observational Data Analysis

Today’s computational, experimental, and observational facilities afford us tremendous opportunities to couple theory and experiment at increasingly large scales. Empirical sensing capabilities are growing dramatically with beam line and detector improvements, and with advances in our ability to deploy large-scale data gathering observations of the natural world. The coupling of computational simulations and analysis to process the data from experimental and observational facilities is giving rise to cross-facility workflows. Such federations of facilities are in fact becoming an explicit requirement for large-scale scientific discovery. As we scale up these pipelines of scientific discovery, each participating facility needs to establish and align policies so that the federation can work seamlessly in an end-to-end manner. This chapter outlines specific policy considerations in enabling the federation of facilities for data analysis. Design choices and vital policy decisions cover the areas of data acquisition and storage, data transfer, computational resource allocation and co-scheduling, seamless federated user access, and cross-cutting governance. By highlighting the explicit and implicit interdependencies between facilities, we aim to provide facility designers and policymakers the information on policy issues to address early in a facility’s operations, thus enabling successful cross-facility federation and improved experimental and observational data analysis outcomes.

Shankar, Mallikarjun (Arjun)↗

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery↗

Position Papers for the ASCR Workshop on the Science of Scientific-Software Development and Use

Software is an increasingly important component in the pursuit of scientific discovery. Both its development and use are essential activities for many scientific teams. At the same time, very little scientific study has been conducted to understand, characterize, and improve the development and use of software for science. Computational science teams have diversified over time to include contributions from domain scientists who provide expertise in scientific and engineering disciplines, applied mathematicians and computer scientists who provide optimal algorithms and data structures, and software and data engineers who provide methodologies and tools adapted and adopted from other software domains. These diverse contributions have enabled tremendous advances in the pursuit of scientific discovery, even as models, computer architectures, and software environments have become more complicated. With this increasing diversity, we believe the next opportunity for qualitative improvement comes from applying the scientific method to understanding, characterizing, and improving how scientific software is developed and used. We believe that this pursuit requires expertise from computational scientists themselves, and from the cognitive and social sciences as well as the software engineering research community. As we look to increase the productivity and sustainability of the scientific-software-development-and-use cycle, a more systematic application of the scientific method to understand processes for software development and use will be a valuable tool to guide future work and result in more usable and sustainable software. This workshop will bring together computer scientists, software engineering researchers, computational scientists, applied mathematicians, social scientists, cognitive scientists, and others, to explore how we can conduct such systematic investigations, what can be learned, and how doing so will benefit the scientific enterprise. The workshop will be structured around a set of breakout sessions, with every attendee expected to participate actively in the discussions. Afterward, workshop attendees — from DOE, industry, and academia — will produce a report for ASCR that summarizes the findings of the workshop.

42 ENGINEERING↗

Environmental Extremes: Building a New Research Partnership between the University of Nevada Reno and the Pacific Northwest National Laboratory (Final Report)

This project focused on strengthening research partnership and collaboration between the University of Nevada, Reno (UNR) and the Pacific Northwest National laboratory (PNNL) in two focused areas: wildfires and hydrology in the context of environmental extremes. The goal was to overcome barriers and bring UNR university researchers, students, and postdocs up to speed and entrained into DOE/SC/BER /Earth and Environmental System Science Division’s (EESSD) environmental research enterprise. UNR continues to work with the Earth and Biological Sciences Directorate (EBSD) at PNNL to create a broad new collaborative research program connecting scientists and research faculty at the two institutions. This research partnership has been designed to increase the capabilities and velocities of both institutions by developing long-term relationships between researchers from each institution, exposing university researchers to the deep capabilities in the DOE National Laboratories and user facilities, and establishing a pipeline of skilled and experienced graduate students and postdoctoral researchers ready to work alongside DOE scientists to take on grand challenge science problems. These collaborations are intended to ultimately benefit the public by advancing scientific discovery, unleashing the scientific potentials through collaboration and connection, enriching the technical and scientific competitiveness of our nation, and enabling the development of a skilled workforce to tackle the critical scientific challenges that face our nation’s security, infrastructure, and prosperity.

54 ENVIRONMENTAL SCIENCES↗

Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Scientific computing is undergoing rapid transformation as advances in artificial intelligence, heterogeneous computing, automation, and data-intensive research reshape not only computational tools but also the institutions, workforce models, and collaborative practices that support scientific discovery. This report synthesizes insights from the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing, the second in a three-year series focused on strengthening scientific computing ecosystems through socio-technical co-design. Workshop discussions identified four interdependent strategic themes: software ecosystems for AI-enabled scientific discovery; trust, validation, and traceability; human-AI teaming and paradigm shifts; and workforce, pedagogy, and governance. The report translates these themes into eight priorities for community action spanning shared research infrastructure, trust and traceability, user experience, human-AI teaming, workforce development, cross-sector coordination, stewardship and sustainability, and evaluation of scientific value. Together, these priorities outline directions for building scientific computing ecosystems that remain trustworthy, sustainable, innovative, and resilient as AI assumes a growing role in scientific work.

AI↗

The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science

Modern scientific discovery increasingly requires coordinating distributed facilities and heterogeneous resources, forcing researchers to act as manual workflow coordinators rather than scientists. Advances in AI leading to AI agents show exciting new opportunities that can accelerate scientific discovery by providing intelligence as a component in the ecosystem. However, it is unclear how this new capability would materialize and integrate in the real world. To address this, we propose a conceptual framework where workflows evolve along two dimensions which are intelligence (from static to intelligent) and composition (from single to swarm) to chart an evolutionary path from current workflow management systems to fully autonomous scientific laboratories. With these trajectories in mind, we present an architectural blueprint that can help the community take the next steps towards harnessing the opportunities in autonomous science with the potential for 100x discovery acceleration and transformational scientific workflows.

Shin, Woong [ORNL] (ORCID:0000000172077814)↗

Toward a Seamless Integration of Computing, Experimental, and Observational Science Facilities: A Blueprint to Accelerate Discovery

The Department of Energy, Office of Science operates world-leading facilities for experimental, observational, and computational science. DOE supercomputing facilities will reach performance at the scale of ExaFLOPs in the coming years, enabling new vistas of scale and precision for large scale simulations and data analysis. Experimental scientific facilities are undergoing similar upgrades that will lead to higher data rates and correspondingly larger computational demands, and will increase the need for near-real-time processing and resilient support for more complex workflows. A transformation of science is underway, with workloads at supercomputing facilities increasingly driven by this explosion of data from instruments and experimental facilities, as well as the accelerating use of Artificial Intelligence (AI) as a tool for scientific discovery. A seamless integration of computing, networking, instruments, and experimental facilities is required to support these emerging workloads and open up a new frontier of U.S. leadership in scientific discovery. We propose to accomplish this by providing frictionless access to the ASCR supercomputing facilities. We describe our vision of combining the power of ASCR supercomputers and networking infrastructure into an integrated scalable fabric, available to end user scientists via interfaces that aim to automate and simplify access to high performance computing systems. This will enable unprecedented computational science capabilities for experimental and observational facilities, and will create new opportunities to combine large simulations and modeling with experimental facility data analysis. This blueprint for creating an integrated network of computational and experimental facilities will provide an enriched discovery environment and open doors for new scientific communities to access the DOE’s world-leading computing and networking capabilities.

97 MATHEMATICS AND COMPUTING↗

Potential Applications of Quantum Computing at Los Alamos National Laboratory, v0.3.0

Since the scientific revolution in the 16th and 17th centuries, the process of scientific discovery has followed an iterative feedback process of observation, hypothesis development and testing with physical experiments, which is widely referred to as the scientific method. This process remained largely unchanged until the middle of the 20th century, when the emergence of digital computers empowered scientist to build and inspect detailed simulations of physical phenomena. Over the last century, computational tools have transformed modern approaches to scientific discovery by enabling fast and affordable hypothesis testing before physical experiments are conducted, shown in Figure 1-1. Some notable examples include: global climate forecasts to understand how the environment may change over decades [130]; modeling the behavior of plasma to design fusion reactors [59]; and understanding the behavior of molecules in biological processes [161, 223].

36 MATERIALS SCIENCE↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Providing Affordable Access to the Lunar and Martian Gravity Environments for Conducting Scientific Research and Enabling Technology Development

NASA’s Artemis campaign aims to explore the Moon for scientific discovery, technology advancement, and to learn how to live and work on another world as we prepare for human missions to Mars. Like the Moon, Mars is a rich destination for scientific discovery and a driver of technologies that will enable humans to travel and explore far from Earth. Conducting space research and validating gravity-dependent technologies in low-Earth orbit or on the surface of the Moon or Mars is expensive and typically requires long project life cycles. To achieve NASA’s Moon to Mars objectives, it will be necessary to have affordable, ground-based test platforms for carrying out this research and technology development in the relevant gravity environment. A proposed upgrade to NASA Glenn Research Center’s Zero Gravity Research Facility is presented that will provide more than 10 seconds of any prescribed gravity level between microgravity and Earth gravity, allowing users to rapidly conduct dozens of tests per day in a microgravity, Lunar, and Martian gravity environment. The design concept is described along with expected performance characteristics, and technologies and research areas that will benefit. The proposed ground-based variable gravity test capability will enable scientific breakthroughs and Moon to Mars technology and subsystem development while also providing a tool for educating and inspiring future scientists and engineers.

Variable Gravity↗

Applications and Techniques for Fast Machine Learning in Science

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science—the concept of integrating powerful ML methods into the real-time experimental data processing loop to accelerate scientific discovery. The material for the report builds on two workshops held by the Fast ML for Science community and covers three main areas: applications for fast ML across a number of scientific domains; techniques for training and implementing performant and resource-efficient ML algorithms; and computing architectures, platforms, and technologies for deploying these algorithms. We also present overlapping challenges across the multiple scientific domains where common solutions can be found. This community report is intended to give plenty of examples and inspiration for scientific discovery through integrated and accelerated ML solutions. This is followed by a high-level overview and organization of technical advances, including an abundance of pointers to source material, which can enable these breakthroughs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

2003 Mars Exploration Rover Mission: Robotic Field Geologists for a Mars Sample Return Mission

The Mars Exploration Rover (MER) Spirit landed in Gusev crater on Jan. 4, 2004 and the rover Opportunity arrived on the plains of Meridiani Planum on Jan. 25, 2004. The rovers continue to return new discoveries after 4 continuous Earth years of operations on the surface of the red planet. Spirit has successfully traversed 7.5 km over the Gusev crater plains, ascended to the top of Husband Hill, and entered into the Inner Basin of the Columbia Hills. Opportunity has traveled nearly 12 km over flat plains of Meridiani and descended into several impact craters. Spirit and Opportunity carry an integrated suite of scientific instruments and tools called the Athena science payload. The Athena science payload consists of the 1) Panoramic Camera (Pancam) that provides high-resolution, color stereo imaging, 2) Miniature Thermal Emission Spectrometer (Mini-TES) that provides spectral cubes at mid-infrared wavelengths, 3) Microscopic Imager (MI) for close-up imaging, 4) Alpha Particle X-Ray Spectrometer (APXS) for elemental chemistry, 5) Moessbauer Spectrometer (MB) for the mineralogy of Fe-bearing materials, 6) Rock Abrasion Tool (RAT) for removing dusty and weathered surfaces and exposing fresh rock underneath, and 7) Magnetic Properties Experiment that allow the instruments to study the composition of magnetic martian materials [1]. The primary objective of the Athena science investigation is to explore two sites on the martian surface where water may once have been present, and to assess past environmental conditions at those sites and their suitability for life. The Athena science instruments have made numerous scientific discoveries over the 4 plus years of operations. The objectives of this paper are to 1) describe the major scientific discoveries of the MER robotic field geologists and 2) briefly summarize what major outstanding questions were not answered by MER that might be addressed by returning samples to our laboratories on Earth.

Ming, Douglas W.↗

Toward digital design at the exascale: An overview of project ICECap

High performance computing has entered the Exascale Age. Capable of performing over 1018 floating point operations per second, exascale computers, such as El Capitan, the National Nuclear Security Administration's first, have the potential to revolutionize the detailed in-depth study of highly complex science and engineering systems. However, in addition to these kind of whole machine “hero” simulations, exascale systems could also enable new paradigms in digital design by making petascale hero runs routine. Currently, untenable problems in complex system design, optimization, model exploration, and scientific discovery could all become possible. Motivated by the challenge of uncovering the next generation of robust high-yield inertial confinement fusion (ICF) designs, project ICECap (Inertial Confinement on El Capitan) attempts to integrate multiple advances in machine learning (ML), scientific workflows, high performance computing, GPU-acceleration, and numerical optimization to prototype such a future. Built on a general framework, ICECap is exploring how these technologies could broadly accelerate scientific discovery on El Capitan. In addition to our requirements, system-level design, and challenges, we describe some of the key technologies in ICECap, including ML replacements for multiphysics packages, tools for human-machine teaming, and algorithms for multifidelity design optimization under uncertainty. As a test of our prototype pre-El Capitan system, we advance the state-of-the art for ICF hohlraum design by demonstrating the optimization of a 17-parameter National Ignition Facility experiment and show that our ML-assisted workflow makes design choices that are consistent with physics intuition, but in an automated, efficient, and mathematically rigorous fashion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

Support for the Core Research Activities and Studies of the Computer Science and Telecommunications Board (DE-SC0020446 Final Technical Report)

Supported the core operations of the National Academies' Computer Science and Telecommunications Board (CSTB). Helped support planning and conducting of board meetings, identification of priority topics in computer science and other areas of computing and communications technologies, and oversight for CSTB's portfolio of studies and convenings. Activities shaped and overseen included: a workshop on Al for scientific discovery; collaborative work with other Academies units on a study on foundational research gaps and future directions for digital twins; a study on current capabilities, future prospects, and governance of facial recognition technologies, a study on post- exascale computing; a study on fostering responsible computing research, collaborative work with other Academies units on automated research workflows for accelerated scientific discovery; a study on meeting federal cybersecurity workforce needs, and a study of the ecosystem driving information technology innovation.

97 MATHEMATICS AND COMPUTING↗

Towards Automated Generation of Chiplet-Based Systems

The Software Defined Architectures (SODA) Synthesizer is an open-source compiler-based tool able to automatically generate domain-specialized systems targeting Application- Specific Integrated Circuits (ASICs) or Field Programmable Gate Arrays (FPGAs) starting from high-level programming. SODA is composed of a high-level frontend, SODA-OPT, which leverages the multilevel intermediate representation (MLIR) framework to interface with productive programming tools (e.g., machine learning frameworks), identify kernels suitable for acceleration, and perform high-level optimizations, and of a state-of-the-art high-level synthesis backend, Bambu from the PandA framework, to generate custom accelerators. One specific application of the SODA Synthesizer is the generation of accelerators to enable ultra-low latency inference and control on autonomous systems for scientific discovery (e.g., electron microscopes, sensors in particle accelerators, etc.). This talk will discuss ongoing work on the SODA synthesizer to enable no-human-in-the-loop generation and design space exploration of the chiplets for highly specialized artificial intelligence accelerators. Connecting these highly specialized chiplets to general-purpose cores or programmable accelerators will allow to quickly deploy autonomous systems for scientific discovery.

Limaye, Ankur M.↗

Does the way we do science foster discovery?

Freedom to explore the unknown is key to scientific discovery. Maximizing modern individualistic measures of scientific productivity like citations and number of publications may impede the progress of science as a whole.

discovery↗

Brochure on the 2024 ASCR Workshop on Energy-Efficient Computing for Science

Large-scale computing has enabled numerous scientific discoveries, including ground-breaking achievements facilitated by the US Department of Energy (DOE) supercomputers and advances in applied mathematics and computer science. While important advances were made in energy efficiency to enable exascale computing, continued efforts are needed to dramatically improve the energy efficiency of the next generation of high-performance computing (HPC) systems and, more broadly, AI data centers. Without substantial improvements in energy efficiency, the energy consumption associated with computing could become a limiting factor for future scientific discovery, national security, and technological advancement.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗