Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “information”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

A physics informed bayesian optimization approach for material design: application to NiTi shape memory alloys

Abstract The design of materials and identification of optimal processing parameters constitute a complex and challenging task, necessitating efficient utilization of available data. Bayesian Optimization (BO) has gained popularity in materials design due to its ability to work with minimal data. However, many BO-based frameworks predominantly rely on statistical information, in the form of input-output data, and assume black-box objective functions. In practice, designers often possess knowledge of the underlying physical laws governing a material system, rendering the objective function not entirely black-box, as some information is partially observable. In this study, we propose a physics-informed BO approach that integrates physics-infused kernels to effectively leverage both statistical and physical information in the decision-making process. We demonstrate that this method significantly improves decision-making efficiency and enables more data-efficient BO. The applicability of this approach is showcased through the design of NiTi shape memory alloys, where the optimal processing parameters are identified to maximize the transformation temperature.

Chemistry↗

Enhancing DESI DR1 full-shape analyses using HOD-informed priors

We present an analysis of DESI Data Release 1 (DR1) that incorporates Halo Occupation Distribution (HOD)-informed priors into Full-Shape (FS) modeling of the power spectrum based on cosmological perturbation theory (PT). By leveraging physical insights from the galaxy-halo connection, these HOD-informed priors on nuisance parameters substantially mitigate projection effects in extended cosmological models that allow for dynamical dark energy. The resulting credible intervals now encompass the posterior maximum from the baseline analysis using gaussian priors, eliminating a significant posterior shift observed in baseline studies. In the ΛCDM framework, a combined DESI DR1 FS information and constraints from the DESI DR1 baryon acoustic oscillations (BAO) — including Big Bang Nucleosynthesis (BBN) constraints and a weak prior on the scalar spectral index — yields Ω m = 0.2994 ± 0.0090 and σ 8 = 0.836$^{+0.024}_{-0.027}$, representing improvements of approximately 4% and 23% over the baseline analysis, respectively. For the w 0 w a CDM model, our results from various data combinations are highly consistent, with all configurations converging to a region with w 0 > -1 and w a < 0. This convergence not only suggests intriguing hints of dynamical dark energy but also underscores the robustness of our HOD-informed prior approach in delivering reliable cosmological constraints.

59 BASIC BIOLOGICAL SCIENCES↗

Probing the limits of cosmological information from the Lyman- α forest 2-point correlation functions

The standard cosmological analysis with the Lyα forest relies on a continuum fitting procedure that suppresses information on large scales and distorts the three-dimensional correlation function on all scales. In this work, we present the first cosmological forecasts without continuum fitting distortion in the Lyα forest, focusing on the recovery of large-scale information. Using idealized synthetic data, we compare the constraining power of the full shape of the Lyα forest auto-correlation and its cross-correlation with quasars using the baseline continuum fitting analysis versus the true continuum. We find that knowledge of the true continuum enables a ∼ 10% reduction in uncertainties on the Alcock-Paczyński (AP) parameter and the matter density, Ω m . We also explore the impact of large-scale information by extending the analysis up to separations of 240 h -1 Mpc along and across the line of sight. The combination of these analysis choices can recover significant large-scale information, yielding up to a ∼ 15% improvement in AP constraints. This improvement is analogous to extending the Lyα forest survey area by ∼ 40%.

Lyman alpha forest↗

Meta-Learning Enhanced Physics-Informed Graph Attention Convolutional Network for Distribution Power System State Estimation

Promptly perceiving distribution system states is challenged by frequent topology changes and uncertain power injections. To address these issues, a Meta-learning enhanced physics-informed graph attention convolutional network (Meta-PIGACN) model is proposed to handle topological variability in distribution system state estimation (DSSE). Specifically, physics information is integrated into the graph convolutional network, enabling a physics-informed edge-weighting process that incorporates physical information to control the aggregation of neighboring nodes. Besides, the graph attention mechanism automatically adjusts the importance of different neighboring nodes, allowing the capture and preservation of inherent system features across varying topologies, thereby improving state estimation accuracy. Furthermore, meta-learning is proposed to acquire empirical knowledge across multiple topologies so that the model can rapidly adapt to new configurations through iterative gradient descent updates even in large-scale systems. In conclusion, the simulation results based on the 33/118/1746-node distribution systems show the high accuracy and efficiency of the proposed model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu↗

Denudation, solute export, landscape evolution modeling, and geographic information system data for the East River watershed, Colorado, USA (2020-2024)

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of lithologic controls on denudation, solute export, carbon-scaling relationships, and transient landscape evolution in the East River watershed near Crested Butte, Colorado, USA. The package includes GIS layers used to produce the Figure 2 map, including drainage, hillshade, lithology, sample locations, and basin polygons, together with comma-separated value (CSV) tables and matching CSV data dictionaries. One group of tables reports sample-level and catchment-level information for river-sediment samples analyzed for in situ-produced cosmogenic beryllium-10 (10Be), including sample names, outlet elevations, geographic coordinates, upstream drainage area, rock-type classes, production-rate scaling scheme, analyzed nuclide, catchment-averaged denudation rates, and associated lower and upper analytical uncertainties. Sample and catchment attributes provide the basis for comparing denudation rates across intrusive, shale, sedimentary, and mixed-lithology settings. A second group of tables reports supporting information for landscape-evolution modeling and the mapped geologic framework of the study area. Included files list parameter values and definitions for the two-phase landscape-evolution simulations, summarize full-domain model erosion fluxes and topographic metrics for different simulation configurations, provide a fixed-area carbon-model scaling table, and summarize mapped geologic units within the East River study domain, including geologic code, formation name, lithologic description, mapped area, and lithologic class grouping. Model outputs and geologic summaries support interpretation of transient landscape behavior and its relation to the mapped distribution of shale, intrusive, sedimentary, and surficial units. A third group of tables reports hydrologic and hydrochemical information used to quantify dissolved export from the watershed. Included files provide site-level values for drainage area, mean annual solute export, standard error of annual export, area-normalized solute yield, and equivalent weathering rate for five East River monitoring sites, along with metadata describing the number, sampling cadence, and date range of discharge records and partial and full total dissolved solids observations used in the solute-yield analyses. The package also contains a supplementary daily ion-load time series with daily mean discharge, discharge observation counts, dissolved concentrations, and daily loads for calcium, magnesium, sodium, potassium, chloride, sulfate, nitrate, fluoride, dissolved silica, charge-balance bicarbonate, and total dissolved solids. The package contains GIS files, comma-separated value files (.csv), CSV data dictionaries, a file-level metadata table, a package-tree text file, and a readme text file.

10Be↗

Radioisotope Identification with List-Mode Gamma Ray Data: A rigorous assessment on the value of temporal information applied to radioisotope identification.

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with “confuser” sources, or spectra with similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research rigorously examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and information theory. We further propose a basic classification model that can utilize spectral or temporal data (or both) to determine if the incorporation of temporal information can improve radioisotope identification. The findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Maximized Information Gain of Next Generation Pulsed Power Using Optimized Design of Z-Machine Experiments

This project develops a Bayesian optimization approach to extracting insights from Z Machine experimental data to determine if and how these insights can be used to extrapolate to a larger facility. The primary goal is to address the scientific challenge of informing how confidently experimental conditions can be predicted on a next generation facility, the design of which requires the reliable extrapolation of current high energy density technologies to regimes yet unobserved, except by costly high-fidelity computational models. Maximizing the use of presently available data and understanding how it informs future endeavors is critically important to enable transformative pulsed power and the science of extreme conditions. We explore a Bayesian optimization approach to experimental design which combines information theory, experimental data, and computational modeling to explore how information gain can be maximized.

97 MATHEMATICS AND COMPUTING↗

Quantum Information Encoding and Decoding for Quantum Sensi

This two-year theory project focused on theoretical investigations of novel paradigms for quantum sensing, building on information encoding and techniques from quantum error correction, quantum computing and other quantum information domains. The outcomes facilitate quantum information technology development, especially at the interface of quantum computing and quantum sensing. The results of the project show new use cases and new paradigms for quantum sensing beyond what has so far been considered. One outcome shows how quantum sensing opens new opportunities for fundamental physics such as the capability of single graviton detection. Another outcome reveals a new application of NISQ quantum computers with error correction for metrology, building on recent advances in practical quantum error correction implementation. The third outcome of the project creates new paradigms of back-action-evading sensing inspired by collective quantum information encoding, which achieves quantum sensing beyond the quantum limit without the use of entanglement.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Detecting flash drought events to inform adaptive water management

Abstract: Flash droughts are characterized by a rapid onset of drought conditions, a common feature among various indicators of these events. The lack of a standardized definition and indicator makes flash droughts particularly challenging to predict and manage. We test six state-of-the-art flash drought indicators using 41 years of daily data across all Hydrological Unit Code 4 (HUC4) regions in the contiguous United States. Our results show substantial disagreements among indicators, suggesting that the detection of a flash drought event strongly depends on the choice of indicator. From the perspective of informing management and operations, an open challenge remains: which flash drought indicator should be used to trigger operational adaptations? To this end, we propose a methodological framework that informs adaptive actions by regionally assessing the level of agreement between different indicators. The regional differences in hydroclimatic conditions are expected to be reflected in the level of agreement between indicators. To better inform adaptive actions, the framework will also consider the impacted sectoral water uses and their varying response times. The insights gained from this study improve the understanding of how these different flash drought indicators capture the different aspects of regional differences and impacted sectors. Furthermore, this information can support more targeted adaptive actions to mitigate adverse impacts through improved preparedness and response strategies.

climate resilience↗

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Stacked networks improve physics-informed training: Applications to neural networks and deep operator networks

Physics-informed neural networks and operator networks have shown promise for effectively solving equations modeling physical systems. However, these networks can happen to be difficult or impossible to train accurately. Here, we present a novel multifidelity framework for stacking physics-informed neural networks and operator networks that facilitates training. We successively build a chain of networks, where the output at one step can act as a low-fidelity input for training a longer chain, gradually increasing the expressivity of the learnt model. The equations imposed at each step of the iterative process can be the same or different (akin to simulated annealing). The iterative (stacking) nature of the proposed method allows us to learn progressively features of a solution which could have been hard to learn directly. Through benchmark problems including a nonlinear pendulum, the wave equation, and the viscous Burgers equation, we show how stacking can be used to improve the accuracy and reduce the required size of physics-informed neural networks and operator networks.

97 MATHEMATICS AND COMPUTING↗

Optimizing the optimizer for physics-informed neural networks and Kolmogorov-Arnold networks

Physics-Informed Neural Networks (PINNs) have revolutionized the computation of PDE solutions by integrating partial differential equations (PDEs) into the neural network’s training process as soft constraints, becoming an important component of the scientific machine learning (SciML) ecosystem. More recently, physics-informed Kolmogorv-Arnold networks (PIKANs) have also shown to be effective and comparable in accuracy with PINNs. In their current implementation, both PINNs and PIKANs are mainly optimized using first-order methods like Adam, as well as quasi-Newton methods such as BFGS and its low-memory variant, L-BFGS. However, these optimizers often struggle with highly nonlinear and non-convex loss landscapes, leading to challenges such as slow convergence, local minima entrapment, and (non)degenerate saddle points. In this study, we investigate the performance of Self- Scaled BFGS (SSBFGS), Self-Scaled Broyden (SSBroyden) methods and other advanced quasi-Newton schemes, including BFGS and L-BFGS with different line search strategies. These methods dynamically rescale updates based on historical gradient information, thus enhancing training efficiency and accuracy. We systematically compare these optimizers – using both PINNs and PIKANs – on key challenging PDEs, including the Burgers, Allen-Cahn, Kuramoto-Sivashinsky, Ginzburg-Landau, and Stokes equations. Additionally, we evaluate the performance of SSBFGS and SSBroyden for Deep Operator Network (DeepONet) architectures, demonstrating their effectiveness for data-driven operator learning. Our findings provide state-of-the-art results with orders-of-magnitude accuracy improvements without the use of adaptive weights or any other enhancements typically employed in PINNs. More broadly, our work reveal insights into the effectiveness of quasi-Newton optimization strategies in significantly improving the convergence and accurate generalization of PINNs and PIKANs.

97 MATHEMATICS AND COMPUTING↗

Modeling information flow in a computer processor with a multi-stage queuing model

In this paper, we introduce a nonlinear stochastic model to describe the propagation of information inside a computer processor. In this model, a computational task is divided into stages, and information can flow from one stage to another. The model is formulated as a spatially-extended, continuous-time Markov chain where space represents different stages. This model is equivalent to a spatially-extended version of the M/M/s queue. The main modeling feature is the throttling function which describes the processor slowdown when the amount of information falls below a certain threshold. We derive the stationary distribution for this stochastic model and develop a closure for a deterministic ODE system that approximates the evolution of the mean and variance of the stochastic model. In conclusion, we demonstrate the validity of the closure with numerical simulations.

97 MATHEMATICS AND COMPUTING↗

Simplifying the Quantum World: Demonstrations for Young Learners in an Informal Setting

A set of modules for the informal learning of quantum science was developed. They include (1) Waves and Bottling Light in Quantum Dots, (2) Quantization of Energy Levels, (3) Particle-Wave Duality, (4) Magnetism and Electron Spin, and (5) Quantum Entanglement. Their teaching objective is to clarify concepts in quantum science, and they have been presented together as part of an hour-long show to ∼250 adults and school-age children. The learning outcomes of the modules were assessed by pre- and postevent quizzes as well as interactive clicker questions. The results suggest effective learning of all of the assessed concepts. These modules are detailed in a way that makes them deployable, together or in part, in other formal or informal settings to support the dissemination of information about quantum science to the general public.

electron spin↗

Scrambling Signal Modularity in Bottom-up Assembled Synthetic Pseudomonas Consortia Reveals Robust Information Transfer

There is immense potential in crafting synthetic microbial communities for application in human health, agriculture, the environment, and even biomanufacturing where an appropriately constructed consortium can be assembled with tremendous biosynthetic or degradative capabilities. In many of these cases, bacterial signaling serves as a form of intercellular information transfer that guides the collective’s behavior. Such communication is complex, as many signals, signal disruptors, microbial species, physical barriers, and spatiotemporal constraints may be involved. Here, in this work, we demonstrate that a multisignal pathway for molecular information transfer within a consortium of several Pseudomonas spp. can be scrambled (genetically and organizationally) while the original message is still effectively conveyed. Assembled from the bottom up, we have employed two types of signaling molecules (i) a redox active secondary metabolite (rhizospheric signal, phloroglucinol), and (ii) a bacterial quorum sensing signal (3-oxo-C12 acylhomoserine lactone, AI-1). These signals can be intraconverted and acted upon by designated community members. We show how the order in which the signals are received, transduced, and subsequently transmitted can be rearranged with minimal impact on the intended outcome. In the consortial context, we found this messaging structure can be remarkably robust. Inspired by rhizospheric molecular signaling mechanisms, this work provides a conceptual framework for designing signaling and information transfer processes within assembled communities.

Biological and medical sciences↗