Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Shift Happens: Building Robust AI Models with Domain Adaptation

Artificial Intelligence (AI) is revolutionizing physics research—from probing the large-scale structure of the Universe to modeling subatomic interactions and fundamental forces. Yet, a major challenge persists: AI models trained on simulations or old experiment / astronomical survey often perform poorly when applied to new data—exposing issues of dataset (domain) shift, model robustness, and uncertainty in predictions. This summer school session will introduce students to common challenges in applying AI across domains and present solutions based on domain adaptation—a set of techniques designed to improve model generalization under domain shift. We will cover foundational ideas, practical strategies, and current research frontiers in this area. Through examples in astrophysics, we'll explore how domain adaptation can help bridge the gap between synthetic and real-world data, improve trust in model outputs, and advance scientific discovery. The concepts discussed are broadly applicable across physics and other scientific disciplines, making this a valuable topic for anyone interested in building robust, transferable AI models for science.

Ciprijanovic, A. [Fermilab] (ORCID:000000031281719↗

High-Throughput Electric-Field-Assisted Sintering and Characterization Techniques for Materials Discovery

Despite improvements in computing and modeling capabilities, the performance of new materials, particularly those which deviate greatly in composition from well-studied materials (e.g., high-entropy alloys), can be difficult to simulate given the lack of available experimental property data. While some modeling techniques may attempt to predict the properties of these exotic materials, most are forced to make extrapolations from more traditional materials. To fulfill the need for accelerated material synthesis and property measurement, a high-throughput methodology has been developed. Utilizing electric-field-assisted sintering (EFAS), also known as spark plasma sintering (SPS), equipped with custom tooling, samples of differing alloy compositions can be produced simultaneously as a single alloy array. Several arrays have been produced with compositions spanning the Co-Cr-Fe-Mn-Ni alloy family, including many high-entropy alloys, while the novel array geometry has enabled the samples to be polished and characterized in parallel, using X-ray diffraction, scanning-electron microscopy, and laser-based thermal diffusivity measurements.

36 MATERIALS SCIENCE↗

Evidence for spreading in the lower Kam Group of the Yellowknife greenstone belt: Implications for Archaean basin evolution in the Slave Province

The Yellowknife greenstone belt is the western margin of an Archean turbidite-filled basin bordered on the east by the Cameron River and Beaulieu River volcanic belts (Henderson, 1981; Lambert, 1982). This model implies that rifting was entirely ensialic and did not proceed beyond the graben stage. Volcanism is assumed to have been restricted to the boundary faults, and the basin was floored by a downfaulted granitic basement. On the other hand, the enormous thickness of submarine volcanic rocks and the presence of a spreading complex at the base of the Kam Group suggest that volcanic rocks were much more widespread than indicated by their present distribution. Rather than resembling volcanic sequences in intracratonic graben structures, the Kam Group and its tectonic setting within the Yellowknife greenstone belt have greater affinities to the Rocas Verdes of southern Chile, Mesozoic ophiolites, that were formed in an arc-related marginal basin setting. The similarities of these ophiolites with some Archean volcanic sequences was previously recognized, and served as basis for their marginal-basin model of greenstone belts. The discovery of a multiple and sheeted dike complex in the Kam Group confirms that features typical of Phanerozoic ophiolites are indeed preserved in some greenstone belts and provides further field evidence in support of such a model.

Helmstaedt, H.↗

Separating Spatial and Temporal Variations of the Aurora Using Two Nearly Colocated Satellites

This final report describes the efforts accomplished during the grant's period of performance, covering the period of 1 May 1997 to 30 April 2001, of a NASA Supporting Research and Technology Program grant under the Ionospheric, Thermospheric, and Mesospheric Physics component of the Sun-Earth Connections program. We have met and exceeded the goals set forth in the proposed research objectives. Referred publications have appeared in the scientific literature and several others are in the review process. In addition, numerous invited and contributed presentations of these studies were presented at national and international meetings during the performance period. One graduate student completed his PhD and won two AGU Best Student Paper awards based on research funded by this grant. These studies are summarized below. The science goal delineated in the initial proposal was "to systematically explore the temporal and spatial characteristics of the aurora in a way heretofore impossible, using data from two coplanar DMSP spacecraft." We accomplished this goal through a series of related studies. One study used these unique data to establish the role of Ps6 waves in coupling between the magnetosphere and the auroral ionosphere (omega bands) during the recovery phase of a magnetic storm; the published paper demonstrated the causal relationships between geospace processes occurring in different regions and established a simple conceptual model based on the fortuitous constellation of observations. In the second string of papers, we used these data to explore velocity-dispersed ions (VDIS) in and near the cusp, to test region identification models, and to look at space/time structure of auroral precipitation. On the first topic, the unique DMSP data revealed a remarkable double VDIS with a latitudinal overlap. This could only be explained in terms of a unified reconnection geometry that builds on several earlier unrelated models. The paper outlining this discovery has drawn considerable attention from the community and is currently in press - it adds significantly to the debate over whether reconnection is study state versus bursty and patchy versus global. The second paper develops the model further by incorporating the electron signature - these ionospheric particle precipitation signatures reveal the presence of magnetospheric "fossilized" FTEs, demonstrating the power of ionospheric measurements as a remote diagnostic of magnetospheric processes. Finally, the general nature of aurora] stability and coherence and region identification by particle characteristics were fully explored in a final paper. We identify candidate mechanisms controlling coherence time scales and length scales and refine boundary region identification criteria. We also use the dual-DMSP observations to identify the open and closed LLBL region and related its significance to the generalized bursty, multiple x-line model developed in the first paper. All of these topics are chapters of Dr. Boudouridis' recently completed PhD thesis.

Spence, Harlan E.↗

Accelerating the discovery of battery electrode materials through data mining and deep learning models

The availability of crystalline materials databases allows for building accurate machine learning (ML) models that can accelerate the exploration of materials chemical space for energy storage applications. In this work, we screen all inorganic materials included in the Materials Project and AFLOW databases as potential metal-ion battery electrodes. We develop an efficient protocol to mine and screen raw data in current databases and provide a new database of electrode materials by considering pairs of charged and discharged electrodes. This effort leads to a new database with over 190,000 instances, in contrast to the original battery database which contains about 5000. The expanded battery data set is then used to build regression-based deep neural network models for predicting average voltages and percentage volume changes upon charging and discharging, which present improvements of at least 28% for target properties with respect to previous models, and are now able to predict anode electrodes (low voltage region) as well as electrodes that will not work in electrochemical cells (negative voltages), overcoming the challenges identified in previous ML models for battery electrodes. Additionally, a further screening of the expanded database itself allowed us to identify 35 novel electrode candidates with excellent battery performance metrics.

25 ENERGY STORAGE↗

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery↗

Sequence-based generative AI design of versatile tryptophan synthases

Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β-subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.

biocatalysis↗

Complexity for Survival of Living Systems

A logical connection between the survivability of living systems and the complexity of their behavior (equivalently, mental complexity) has been established. This connection is an important intermediate result of continuing research on mathematical models that could constitute a unified representation of the evolution of both living and non-living systems. Earlier results of this research were reported in several prior NASA Tech Briefs articles, the two most relevant being Characteristics of Dynamics of Intelligent Systems (NPO- 21037), NASA Tech Briefs, Vol. 26, No. 12 (December 2002), page 48; and Self-Supervised Dynamical Systems (NPO- 30634) NASA Tech Briefs, Vol. 27, No. 3 (March 2003), page 72. As used here, living systems is synonymous with active systems and intelligent systems. The quoted terms can signify artificial agents (e.g., suitably programmed computers) or natural biological systems ranging from single-cell organisms at one extreme to the whole of human society at the other extreme. One of the requirements that must be satisfied in mathematical modeling of living systems is reconciliation of evolution of life with the second law of thermodynamics. In the approach followed in this research, this reconciliation is effected by means of a model, inspired partly by quantum mechanics, in which the quantum potential is replaced with an information potential. The model captures the most fundamental property of life - the ability to evolve from disorder to order without any external interference. The model incorporates the equations of classical dynamics, including Newton s equations of motion and equations for random components caused by uncertainties in initial conditions and by Langevin forces. The equations of classical dynamics are coupled with corresponding Liouville or Fokker-Planck equations that describe the evolutions of probability densities that represent the uncertainties. The coupling is effected by fictitious information-based forces that are gradients of the information potential, which, in turn, is a function of the probability densities. The probability densities are associated with mental images both self-image and nonself images (images of external objects that can include other agents). The evolution of the probability densities represents mental dynamics. Then the interaction between the physical and metal aspects of behavior is implemented by feedback from mental to motor dynamics, as represented by the aforementioned fictitious forces. The interaction of a system with its self and nonself images affords unlimited capacity for increase of complexity. There is a biological basis for this model of mental dynamics in the discovery of mirror neurons that learn by imitation. The levels of complexity attained by use of this model match those observed in living systems. To establish a mechanism for increasing the complexity of dynamics of an active system, the model enables exploitation of a chain of reflections exemplified by questions of the form, "What do you think that I think that you think...?" Mathematically, each level of reflection is represented in the form of an attractor performing the corresponding level of abstraction with more details removed from higher levels. The model can be used to describe the behaviors, not only of biological systems, but also of ecological, social, and economics ones.

Zak, Michail↗

Exploring cryo-electron microscopy with molecular dynamics

Single particle analysis cryo-electron microscopy (EM) and molecular dynamics (MD) have been complimentary methods since cryo-EM was first applied to the field of structural biology. The relationship started by biasing structural models to fit low-resolution cryo-EM maps of large macromolecular complexes not amenable to crystallization. The connection between cryo-EM and MD evolved as cryo-EM maps improved in resolution, allowing advanced sampling algorithms to simultaneously refine backbone and sidechains. Moving beyond a single static snapshot, modern inferencing approaches integrate cryo-EM and MD to generate structural ensembles from cryo-EM map data or directly from the particle images themselves. We summarize the recent history of MD innovations in the area of cryo-EM modeling. The merits for the myriad of MD based cryo-EM modeling methods are discussed, as well as, the discoveries that were made possible by the integration of molecular modeling with cryo-EM. Lastly, current challenges and potential opportunities are reviewed.

Biochemistry & Molecular Biology↗

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery↗

Superdiffusion resilience in Heisenberg chains with two-dimensional interactions on a quantum processor

Superdiffusive spin transport in the one-dimensional (1D) Heisenberg model is a key theoretical discovery in nonequilibrium quantum many-body physics. Although extensively studied in 1D systems, the breakdown and sustenance of superdiffusion in two-dimensional (2D) lattices with integrability-breaking terms, as found in real materials, remains an open question. To address this, we develop a toy model that extends the 1D Heisenberg model with a representative set of 2D interaction types and tunable strengths. Our model exhibits varying degrees of superdiffusion breakdown depending on the interaction type, spanning ballistic to diffusive regimes. We establish and justify a hierarchy of 2D interactions based on their resilience against superdiffusion breakdown: Heisenberg >𝑋⁢𝑋 > Ising. This precise control over the superdiffusive behavior also enables rigorous benchmarking of quantum hardware, and our simulations on IBM's Heron devices confirm the hardware's ability to accurately capture these many-body nonequilibrium phenomena. Overall, our results are relevant not only to simulating superdiffusion in real materials, such as the 1D Heisenberg compound KCuF3, which contains modest nonintegrable 2D terms, but also to extending superdiffusive behavior to larger 2D qubit lattices and other 2D materials.

Alagarsamy Manikandan, Keerthi Kumaran [ORNL]↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

The Thermococcales as a model system: historical perspectives and emerging tools

Thermococcales are among the most widely studied hyperthermophilic Archaea and have become key models for understanding life at extreme temperatures. Early work in the 1980s culminated in the isolation of novel Thermococcales species from hydrothermal vents that grew rapidly, tolerated extreme heat, and metabolized diverse substrates, making them uniquely amenable for laboratory studies. Their thermostable enzymes and emerging genetic tools facilitated detailed investigations of core processes such as DNA replication, repair, and transcription under conditions that challenge most life forms. These practical advantages, together with the accumulation of tools and protocols, cemented the role of Thermococcales as a model system. Here, we recount how chance discoveries, environmental adaptations, and experimental practicality intersected to establish Thermococcales as a central model for studying archaeal biology and extremophile physiology.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning techniques for intermediate mass gap lepton partner searches at the large hadron collider

We consider machine learning techniques associated with the application of a boosted decision tree (BDT) to searches at the Large Hadron Collider (LHC) for pair-produced lepton partners which decay to leptons and invisible particles. This scenario can arise in the minimal supersymmetric Standard Model (MSSM), but can be realized in many other extensions of the Standard Model (SM). We focus on the case of intermediate mass splitting ( ∼ 30 GeV ) between the dark matter (DM) and the scalar. For these mass splittings, the LHC has made little improvement over LEP due to large electroweak backgrounds. We find that the use of machine learning techniques can push the LHC well past discovery sensitivity for a benchmark model with a lepton partner mass of ∼ 110 GeV , for an integrated luminosity of 300 fb − 1 , with a signal-to-background ratio of ∼ 0.3 . The LHC could exclude models with a lepton partner mass as large as ∼ 160 GeV with the same luminosity. The use of machine learning techniques in searches for scalar lepton partners at the LHC could thus definitively probe the parameter space of the MSSM in which scalar muon mediated interactions between SM muons and Majorana singlet DM can both deplete the relic density through dark matter annihilation and satisfy the recently measured anomalous magnetic moment of the muon. We identify several machine learning techniques which can be useful in other LHC searches involving large and complex backgrounds. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

The tree of life describes a tripartite cellular world

The canonical view of a 3-domain (3D) tree of life was recently challenged by the discovery of Asgardarchaeota encoding eukaryote signature proteins (ESPs), which were treated as missing links of a 2-domain (2D) tree. Here we revisit the debate. We discuss methodological limitations of building trees with alignment-dependent approaches, which often fail to satisfactorily address the problem of ‘‘gaps.’’ In addition, most phylogenies are reconstructed unrooted, neglecting the power of direct rooting methods. Alignment-free methodologies lift most difficulties but require employing realistic evolutionary models. We argue that the discoveries of Asgards and ESPs, by themselves, do not rule out the 3D tree, which is strongly supported by comparative and evolutionary genomic analyses and vast genomic and biochemical superkingdom distinctions. Given uncertainties of retrodiction and interpretation difficulties, we conclude that the 3D view has not been falsified but instead has been strengthened by genomic analyses. In turn, the objections to the 2D model have not been lifted. The debate remains open. Also see the video abstract here: https://youtu.be/-6TBN0bubI8

59 BASIC BIOLOGICAL SCIENCES↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Chemical Blast Standard (1 kg)

Chemical explosions create blast waves with large overpressure disturbances. It is important to develop a standard blast model based on data to accurately predict acoustic blast-wave amplitudes near detonations and invert for explosion energy from distant observations of blast-wave signals. However, open data from large, controlled chemical explosions with reliable ground truth can be challenging to find. The lack of access to such data could limit the number of contributions to related research and potentially stifle the rate of discoveries or validation of existing models. Here, to address these data scarcity problem, we have curated and compiled a standardized set of 817 blast-wave waveforms from 19 distinct high-explosive events. The blast-wave waveforms are standardized to a 1 kg trinitrotoluene explosion using scaling laws and corrections for location effects. A brief overview of the dataset is presented along with explosion feature models as well as recommendations for extracting explosion features. The resulting dataset is distributed to an open repository in both Seismic Analysis Code and pandas DataFrame formats containing the waveforms, the scaled distances, and the sample rates.

58 GEOSCIENCES↗