Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hypothesis learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Co-optimization of fuel properties, combustion system geometry, and injection strategy for conventional diesel fuel

Here, studies have shown that fuel properties can impact an engine’s operation in several ways, including ignition delay, sooting tendency, mixture formation, and combustion temperature. In mixing-controlled compression ignition (MCCI) engines, the fuel system design and piston bowl geometry significantly affect combustion performance and emissions. Based on current information, it is difficult to draw conclusions about fuel property effects and sensitivities. The central fuel hypothesis approach used in the US Department of Energy Co-Optima program has worked well for spark ignition fuels: identifying critical fuel property ranges is sufficient to screen fuel blends that are expected to maximize efficiency and reduce pollutant emissions. However, for MCCI-relevant fuels, the information gained from past studies is not sufficient to build such a merit function or to allow for performing a similar screening of fuel blends. It is hypothesized that a co-optimization of a fuel’s physical and chemical properties, combustion system geometry, and injection strategy could leverage synergies between the effects of the fuel properties and geometries, resulting in improved performance over state-of-the-art. A machine learning–assisted unconstrained global optimization algorithm was used to explore a design space comprising 23 independent variables. The results show that physical property effects were minimal even for large variations in fuel properties, and the only interaction effect that was observed was the effect of varied fuel density parameters on fuel/air mixture formation. Nevertheless, these interactions were not sufficient in magnitude to significantly affect optimization results. Therefore, analysis of the results suggests that fuel physical properties cannot be leveraged in a co-optimization context to increase engine efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Zentropy Theory for Transformative Functionalities of Magnetic and Superconducting Materials

The proposed research developed the zentropy theory through applications to complex magnetic materials and superconductors under the hypothesis that the emergent properties of complex magnetic materials and superconductors can be predicted by statistical mechanics of ergodic microstates with their partition functions computed from DFT-predicted free energies. The key objective is to develop approaches to systematically determine the types and number of microstates and the supercell size in DFT-based calculations through convergency of macroscopic functionalities, with the incorporation of our mixed-space approach accounting for the interactions between periodic supercells. In addition to use scientific intuitions to guide the design of important microstates, the key innovation of the proposed research is to integrate the domain knowledge and the material-property-descriptor database (MPDD) with 4 million microstates, which is supported by our deep neural network machine learning models (SIPFENN: structure-informed prediction of formation energy using neural networks) and integrated with our high throughput DFT Tool Kit (DFTTK). For complex magnetic materials, one of the objectives is to develop approaches to calculate short-range ordering from the statistical distribution of each microstate. For superconductors, the divergency of quasiparticle effective mass at a quantum critical point will be investigated, and the superconducting and non-superconducting microstates will be delineated through analysis of electronic band structure, density of states, charge density, and Fermi surface.

36 MATERIALS SCIENCE↗

Data-Driven RANS Turbulence Closures for Forced Convection Flow in Reactor Downcomer Geometry

Recent progress in data-driven turbulence modeling has shown its potential to enhance or replace traditional equation-based Reynolds-averaged Navier-Stokes (RANS) turbulence models. Here, this work utilizes invariant neural network (NN) architectures to model Reynolds stresses and turbulent heat fluxes in forced convection flows (when the models can be decoupled). As the considered flow is statistically one dimensional, the invariant NN architecture for the Reynolds stress model reduces to the linear eddy viscosity model. To develop the data-driven models, direct numerical and RANS simulations in vertical planar channel geometry mimicking a part of the reactor downcomer are performed. Different conditions and fluids relevant to advanced reactors (sodium, lead, unitary-Prandtl-number fluid, and molten salt) constitute the training database. The models enabled accurate predictions of velocity and temperature, and compared to the baseline k–τ turbulence model with the simple gradient diffusion hypothesis, do not require tuning of the turbulent Prandtl number. The data-driven framework is implemented in the open-source graphics processing unit–accelerated spectral element solver nekRS and has shown the potential for future developments and consideration of more complex mixed convection flows.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Ecohydrological controls on root and microbial respiration in the East River watershed of Colorado

The main objective of this project was to conduct exploratory work to quantify how snow and rain water inputs influence the CO2 coming from the soil surface (soil CO2 flux), and its plant and microbial sources, in the East River watershed, near Crested Butte, Colorado. Knowledge gained from this effort laid the groundwork for a more comprehensive (ongoing) follow-on grant that is using a combination of experiments, field observations, machine learning and simulation modeling to fully disentangle these relationships. New field measurements were made at four locations along an elevational transect on Snodgrass Mountain at Rocky Mountain Biological Laboratory. These sites were chosen to differ in total snowpack and water table depth, and included both deciduous (aspen) and evergreen (spruce/fir) forest types. We used automated measurements of soil CO2 concentrations to quantify the total soil CO2 flux, and the vertical CO2 production within the soil profile at each site. Isotope (radiocarbon, 14C) measurements determined how much of the CO2 emitted from the soil surface came from plant respiration (root metabolism) versus microbial respiration (decomposition of soil organic matter) sources. Supporting data on plant phenology and microbial dynamics provided context for the observed variation in respiration sources. This work was motivated by our overarching hypothesis that quantifying belowground plant and microbial processes separately, and how they are influenced by snow and rain inputs, is necessary for understanding and predicting how the East River watershed ecosystems will respond to future environmental change.

54 ENVIRONMENTAL SCIENCES↗

Search for CP violation in events with top quarks and Z bosons at $\sqrt{s}$ = 13 and 13.6 TeV

A search for the violation of the charge-parity (CP) symmetry in the production of top quarks in association with Z bosons is presented, using events with at least three charged leptons and additional jets. The search is performed in a sample of proton-proton collision data collected by the CMS experiment at the CERN LHC in 2016–2018 at a center-of-mass energy of 13 TeV and in 2022 at 13.6 TeV, corresponding to a total integrated luminosity of 173 fb –1 . For the first time in this final state, observables that are odd under the CP transformation are employed. Also for the first time, physics-informed machine-learning techniques are used to construct these observables. While for standard model (SM) processes the distributions of these observables are predicted to be symmetric around zero, CP-violating modifications of the SM would introduce asymmetries. Two CP-odd operators $\mathcal{O}$$^{I}_{tW}$ and $\mathcal{O}$$^{I}_{tZ}$ in the SM effective field theory are considered that may modify the interactions between top quarks and electroweak bosons. The obtained results are consistent with the SM prediction within two standard deviations, and exclusion limits on the associated Wilson coefficients of –2.7 < $c$$^{I}_{tW}$ < 2.5 and –0.2 < $c$$^{I}_{tZ}$ < 2.0 and are set at 95 % confidence level. The largest discrepancy is observed in $c$$^{I}_{tZ}$ where data is consistent with positive values, with an observed local significance with respect to the SM hypothesis of 2.5 standard deviations, when only linear terms are considered.

CMS↗

Process Anomaly Detection for Sparsely Labeled Events in Nuclear Power Plants

An essential aspect of online monitoring, subtle anomaly detection increases the detection lead time for equipment failure and enables a nuclear power plant (NPP) to mitigate unexpected partial or full outages, resulting in significant cost saving to the plant. Once an anomaly is detected by plant staff, its cause and severity are investigated. Because the vast majority of anomalies require some level of investigation, including some that require time-consuming examination, before they are passed over to the engineering organization for further analysis, plants are often equipped with tools to assist the staff in performing anomaly detection. Those tools operate as a black box and are often based on statistical methods that establish sensor correlations using preconfigured mathematical models and flag correlation deviations as anomalies. Due to the number of anomalies detected at a given NPP on a daily basis, a significant number of flagged anomalies usually await examination for days or weeks. A primary cause of this backlog is that the methods used by the tools generate many false positives. Though this is usually attributed to oversensitive model settings due to very narrow normal operation bands, it can also be associated with the model development being inadequate for the process being monitored, or with missing model inputs that could have explained misclassified positives. The performance of anomaly detection tools impacts their plant acceptance and utilization, especially when the effort to address false positives generated by the tool depletes the value or cost saved by using that tool. Thus, means to advance anomaly detection performance have been investigated by the Department of Energy’s Light Water Reactor Sustainability program. Previous and ongoing efforts have targeted unsupervised machine-learning (ML) methods, which do not require the labeling of any data fed into the ML model. By contrast, in supervised anomaly detection methods, every data point is labeled as either a normal or abnormal process condition, and the model is trained to replicate the classification process. Supervised methods usually outperform unsupervised methods, due to the added value in differentiating normal from anomalous states of the monitored process. An NPP’s corrective action program requires it to track and document, via a dedicated report, the resolution of any issues that occur within the plant. Once created, each report is reviewed by a plant screening committee, and several classifications and decisions are made. Recently, a collaborating NPP developed an artificial intelligence and ML-based classifier to categorize a condition report (CR) into classes that can serve to label the data as normal or anomalous. Applying CRs as labels represents a semi-supervised use case. Semi-supervised ML assumes that labels exist for some data points (i.e., labeled anomalies, in this case) but not for the rest. In this effort, semi-supervised ML methods were used to fuse data from CRs with anomaly detection methods in order to test the hypothesis that partially labeled anomalies would improve the accuracy of the anomaly detection methods. Specifically, two methods were used. The first is the deep Semi-supervised Anomaly Detection (deep SAD) method, which can handle labels ranging from fully unsupervised to fully supervised cases. The second is a newly designed ML method developed specifically for this effort and referred to as the high-order feature (HOF)-based method. To evaluate these two methods in controlled environments, synthetic data generators were developed and used. The first datasets used a spring-mass-damper (SMD) system simulator commonly found in mechanical engineering references. This was used to create two use cases: a one- and a three-mass system. Anomalies were introduced by changing the spring and damper coefficients while the system was actuated by random forces. The second datasets used the commercial Dymola-Modelica software to build a simplified nuclear reactor model. Anomalies were added in the form of corrupted sensor readings and/or control commands. The deep SAD method was tested using the SMD system, while the HOF method was tested using both datasets. Application of the deep SAD semi-supervised ML method demonstrated that labels can generate increased confidence in detecting true anomalies. This helped increase the number of true positives and decrease the number of false negatives—something that would aid in addressing the backlog of possible anomalies. Application of the HOF method demonstrated that labels can aid in down selecting from a candidate set of features to a more optimal subset in order to better differentiate between normal and anomalous conditions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

The future of self-driving laboratories: from human in the loop interactive AI to gamification

Recent developments in artificial intelligence (AI) and machine learning (ML), implemented through self-driving laboratories (SDLs), are rapidly creating unprecedented opportunities for the accelerated discovery and optimization of materials. This paper provides a joint analysis of SDLs from both academic and industry perspectives, highlighting the importance of integrating human intelligence in these systems. It discusses the necessity of careful planning in SDL design across physical, data, and workflow dimensions, including instrumental setup, experimental workflow, data management, and human–SDL interaction. The significance of integrating human input within SDLs, especially as the focus shifts from individual tools and tasks to the creation and management of complex workflows, is emphasized. The paper stresses on the crucial role of reward function design in developing forward-looking workflows and examines the interplay between hardware evolution, ML application across chemical processes, and the influence of reward systems in research. Ultimately, the article advocates for a future where SDLs blend human intuition in hypothesis formulation with AI's precision, speed, and data-handling capabilities.

97 MATHEMATICS AND COMPUTING↗

Security-by-Design: Light Water Small Modular Reactor

The growing demand for nuclear power is increasing pressure to find solutions to cost prohibitive requirements of both construction and security. Offsite response has been proposed as an option to reduce costs associated with training and maintaining an onsite response force. A previous report explored this option and revealed that security could be provided at the required level, but cost savings was not a result of this methodology. An offsite response strategy required costly active and passive delay barriers to provide sufficient time for responders to muster and deploy to a site in time to interrupt a determined and well-equipped adversary. Also, contrary to the hypothesis, the number of responders required for this strategy exceeded that needed for an onsite response force, as the adversaries could avail themselves of advantageous positions within the facility to repel arriving responders. This report builds upon the previous evaluation by using the same hypothetical light water small modular reactor (LWSMR) facility model, but this time an onsite response strategy was assessed. The goal of this analysis was to show that an onsite response strategy could be implemented effectively at a cost point that removes barriers within the industry at this critical time of growth and development. The assessment of the facility design and response strategy was completed through modeling using Scribe3D© and subsequent scenario analysis over the course of a two-day tabletop exercise. Subject matter experts in nuclear security, nuclear facility design, and response strategy and tactics contributed to the effort to ensure accurate representation of hypothetical scenarios. Several adaptations were made to the layout of the LWSMR based on lessons learned during the first day of scenario analysis. The subsequent design evaluated on the second day proved to provide a robust response posture against a large and well-trained adversary force. This report details the process of the analysis and compares the cost of the final facility design with that of the LWSMR model used for evaluation of offsite response. Ultimately, the results of this effort indicate that, when implemented correctly, an onsite response strategy is the best option from a security and cost perspective.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multidimensional perspectives of geo-epidemiology: from interdisciplinary learning and research to cost–benefit oriented decision-making

Research typically promotes two types of outcomes (inventions and discoveries), which induce a virtuous cycle: something suspected or desired (not previously demonstrated) may become known or feasible once a new tool or procedure is invented and, later, the use of this invention may discover new knowledge. Research also promotes the opposite sequence—from new knowledge to new inventions. This bidirectional process is observed in geo-referenced epidemiology—a field that relates to but may also differ from spatial epidemiology. Geo-epidemiology encompasses several theories and technologies that promote inter/transdisciplinary knowledge integration, education, and research in population health. Based on visual examples derived from geo-referenced studies on epidemics and epizootics, this report demonstrates that this field may extract more (geographically related) information than simple spatial analyses, which then supports more effective and/or less costly interventions. Actual (not simulated) bio-geo-temporal interactions (never captured before the emergence of technologies that analyze geo-referenced data, such as geographical information systems) can now address research questions that relate to several fields, such as Network Theory. Thus, a new opportunity arises before us, which exceeds research: it also demands knowledge integration across disciplines as well as novel educational programs which, to be biomedically and socially justified, should demonstrate cost-effectiveness. Grounded on many bio-temporal-georeferenced examples, this report reviews the literature that supports this hypothesis: novel educational programs that focus on geo-referenced epidemic data may help generate cost-effective policies that prevent or control disease dissemination.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

Synchronization of Alternative Models in a Supermodel and the Learning of Critical Behavior

“Supermodeling” climate by allowing different models to assimilate data from one another in run time has been shown to give results superior to those of any one model and superior to any weighted average of model outputs. The only free parameters, connection strengths between corresponding variables in each pair of models, are determined using some form of machine learning. It is demonstrated that supermodeling succeeds because near critical states, interscale interactions are important but unresolved processes cannot be effectively represented diagnostically in any single parameterization scheme. In two examples, a pair of toy quasigeostrophic (QG) channel models of the midlatitudes and a pair of ECHAM5 models of the tropical Pacific atmosphere with a common ocean, supermodels dynamically combine parameterization schemes so as to capture criticality, associated critical structures, and the supporting scale interactions. The QG supermodeling scheme extends a previous configuration in which two such models synchronize with intermodel connections only between medium-scale components of the flow; here the connections are trained against a third “real” model. Intermittent blocking patterns characterize the critical behavior thus obtained, even where such patterns are missing in the constituent models. In the ECHAM-based climate supermodel, the corresponding critical structure is the single ITCZ pattern, a pattern that occurs in neither of the constituent models. In conclusion, for supermodels of both types, power spectra indicate enhanced interscale interactions in frequency or energy ranges of physical interest, in agreement with observed data, and supporting a generalized form of the self-organized criticality hypothesis.

54 ENVIRONMENTAL SCIENCES↗

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗

Machine Intelligence-Centered System for Automated Characterization of Functional Materials and Interfaces

Classic design of experiment relies on a time-intensive workflow that requires planning, data interpretation, and hypothesis building by experienced researchers. Here, in this paper, we describe an integrated, machine-intelligent experimental system which enables simultaneous dynamic tests of electrical, optical, gravimetric, and viscoelastic properties of materials under a programmable dynamic environment. Specially designed software controls the experiment and performs on-the-fly extensive data analysis and dynamic modeling, real-time iterative feedback for dynamic control of experimental conditions, and rapid visualization of experimental results. The system operates with minimal human intervention and enables time-efficient characterization of complex dynamic multifunctional environmental responses of materials with simultaneous data processing and analytics. The system provides a viable platform for artificial intelligence (AI)-centered material characterization, which, when coupled with an AI-controlled synthesis system, could lead to accelerated discovery of multifunctional materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computational Imaging for Intelligence in Highly Scattering Aerosols (Final Report)

Natural and man-made degraded visual environments pose major threats to national security. The random scattering and absorption of light by tiny particles suspended in the air reduces situational awareness and causes unacceptable down-time for critical systems and operations. To improve the situation, we have developed several approaches to interpret the information contained within scattered light to enhance sensing and imaging in scattering media. These approaches were tested at the Sandia National Laboratory Fog Chamber facility and with tabletop fog chambers. Computationally efficient light transport models were developed and leveraged for computational sensing. The models are based on a weak angular dependence approximation to the Boltzmann or radiative transfer equation that appears to be applicable in both the moderate and highly scattering regimes. After the new model was experimentally validated, statistical approaches for detection, localization, and imaging of objects hidden in fog were developed and demonstrated. A binary hypothesis test and the Neyman-Pearson lemma provided the highest theoretically possible probability of detection for a specified false alarm rate and signal-to-noise ratio. Maximum likelihood estimation allowed estimation of the fog optical properties as well as the position, size, and reflection coefficient of an object in fog. A computational dehazing approach was implemented to reduce the effects of scatter on images, making object features more readily discernible. We have developed, characterized, and deployed a new Tabletop Fog Chamber capable of repeatably generating multiple unique fog-analogues for optical testing in degraded visual environments. We characterized this chamber using both optical and microphysical techniques. In doing so we have explored the ability of droplet nucleation theory to describe the aerosols generated within the chamber, as well as Mie scattering theory to describe the attenuation of light by said aerosols, and correlated the aerosol microphysics to optical properties such as transmission and meteorological optical range (MOR). This chamber has proved highly valuable and has supported multiple efforts inclusive to and exclusive of this LDRD project to test optics in degraded visual environments. Circularly polarized light has been found to maintain its polarization state better than linearly polarized light when propagating through fog. This was demonstrated experimentally in both the visible and short-wave infrared (SWIR) by imaging targets made of different commercially available retroreflective films. It was found that active circularly polarized imaging can increase contrast and range compared to linearly polarized imaging. We have completed an initial investigation of the capability for machine learning methods to reduce the effects of light scattering when imaging through fog. Previously acquired experimental long-wave images were used to train an autoencoder denoising architecture. Overfitting was found to be a problem because of lack of variability in the object type in this data set. The lessons learned were used to collect a well labeled dataset with much more variability using the Tabletop Fog Chamber that will be available for future studies. We have developed several new sensing methods using speckle intensity correlations. First, the ability to image moving objects in fog was shown, establishing that our unique speckle imaging method can be implemented in dynamic scattering media. Second, the speckle decorrelation over time was found to be sensitive to fog composition, implying extensions to fog characterization. Third, the ability to distinguish macroscopically identical objects on a far-subwavelength scale was demonstrated, suggesting numerous applications ranging from nanoscale defect detection to security. Fourth, we have shown the capability to simultaneously image and localize hidden objects, allowing the speckle imaging method to be effective without prior object positional information. Finally, an interferometric effect was presented that illustrates a new approach for analyzing speckle intensity correlations that may lead to more effective ways to localize and image moving objects. All of these results represent significant developments that challenge the limits of the application of speckle imaging and open important application spaces. A theory was developed and simulations were performed to assess the potential transverse resolution benefit of relative motion in structured illumination for radar systems. Results for a simplified radar system model indicate that significant resolution benefits are possible using data from scanning a structured beam over the target, with the use of appropriate signal processing.

58 GEOSCIENCES↗

Heat transport in liquid water from first-principles and deep neural network simulations

In this work, we compute the thermal conductivity of water within linear response theory from equilibrium molecular dynamics simulations, by adopting two different approaches. In one, the potential energy surface (PES) is derived on the fly from the electronic ground state of density functional theory (DFT) and the corresponding analytical expression is used for the energy flux. In the other, the PES is represented by a deep neural network (DNN) trained on DFT data, whereby the PES has an explicit local decomposition and the energy flux takes a particularly simple expression. By virtue of a gauge invariance principle, established by Marcolongo, Umari, and Baroni, the two approaches should be equivalent if the PES were reproduced accurately by the DNN model. We test this hypothesis by calculating the thermal conductivity, at the GGA (PBE) level of theory, using the direct formulation and its DNN proxy, finding that both approaches yield the same conductivity, in excess of the experimental value by approximately 60%. Besides being numerically much more efficient than its direct DFT counterpart, the DNN scheme has the advantage of being easily applicable to more sophisticated DFT approximations, such as meta-GGA and hybrid functionals, for which it would be hard to derive analytically the expression of the energy flux. We find in this way that a DNN model, trained on meta-GGA (SCAN) data, reduces the deviation from experiment of the predicted thermal conductivity by about 50%, leaving the question open as to whether the residual error is due to deficiencies of the functional, to a neglect of nuclear quantum effects in the atomic dynamics, or, likely, to a combination of the two.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

3P Program: Phenotyping X Prediction = Productivity (Final Scientific/Technical Report)

The goal of the 3P Program was to establish integrated, real-time phenotyping and to analyze above- and below-ground plant architecture and total carbon partitioning and allocation to predict heterosis and develop superior crop hybrids by fully leveraging the Sorghum gene pool. There were two overarching themes: 1) the development of a new crop improvement approach utilizing advances in high-throughput phenotyping (HTP), computing, and genomics for public dissemination and 2) leveraging this platform for sorghum crop improvement and commercialization. The Clemson team worked on creating genomic resources and using both statistical learning and high-throughput phenotyping in genomics-assisted breeding. Research was broadly interested in the genetics of carbon partitioning, with the aim of improving crop performance and achieving sustainability. The technology and resources created can be readily found in the public domain and serve to advance scientific understanding of crop genomics and breeding. Genomic prediction was able to identify top crosses to be made, and a hybrid prediction pipeline is in place to drive year-over-year genetic gain. Roots have long been ignored by plant breeders and agronomists, not because they are unimportant but because they are hard to measure. This is an untapped white space of potential insight and innovation. To address this, Hi Fidelity Genetics developed the RootTracker to measure roots in the field on a continuous basis. A database system called RootTracker Tracker was developed to handle data coming from the RootTrackers. In using this device, valuable data was observed for plant breeding, hydrochemical development, and other agricultural biology applications. Carnegie Mellon’s goal was developing new techniques to generate high-resolution 3D models of plants from data collected in the field. The idea was that more useful and more informative phenotypes could be extracted by resolving small features, such as seeds and flowers, and that by modeling in 3D, the spatial structure of plants could be examined. To achieve this, multiple images collected by a new small format structured light stereo imager were fused together. A sorghum panicle modeling pipeline was developed to allow the collection and processing of data. Carolina Seed Systems is an agricultural technology company focused on decarbonizing the agricultural system. Their technology pipeline serves to drive fundamental progress towards creation and distribution of carbon negative crops. The genomic and the engineering technology developed through the 3P Program was leveraged to deliver both value and sustainability from the grower to the consumer. Promising sorghum hybrids were scaled up and commercialized. The overall goal of our research was to integrate, create, and deploy genetic and engineering concepts and technologies to enhance crop productivity in a sustainable fashion. The combination of public and private partners allowed the basic research and hypothesis testing to be quickly accelerated for commercial application by the companies yet maintained that the core framework and academic insights remain in the public domain for continued market disruption, competition, and innovation.

59 BASIC BIOLOGICAL SCIENCES↗

Artificial Intelligence and Machine Learning for Bioenergy Research: Opportunities and Challenges

The integration of artificial intelligence and machine learning (AI/ML) with automated experimentation, genomics, biosystems design, and bioprocessing technologies is poised to revolutionize scientific investigation and, particularly, bioenergy research. To identify the opportunities and challenges in this emerging research area, the U.S. Department of Energy’s (DOE) Biological and Environmental Research program (BER) and Bioenergy Technologies Office (BETO) held a joint virtual workshop on AI/ML for Bioenergy Research (AMBER) on August 23–25, 2022. These interests have since been amplified in a September 2022 Executive Order, “Advancing Biotechnology and Biomanufacturing Innovation for a Sustainable, Safe, and Secure U.S. Bioeconomy,” to promote a whole-of government approach to biotechnology development (White House 2022). Approximately 50 scientists with various backgrounds and expertise from academia, industry, and DOE national laboratories met to discuss the opportunities and challenges of AI/ML for bioenergy research. Workshop participants were tasked with assessing the potential for AI/ML and laboratory automation to advance biological understanding and engineering in general. They particularly examined how integrating AI/ML tools with laboratory automation could accelerate biosystems design and optimize biomanufacturing. Discussions included the data and computational infrastructure needed to augment biosystems design applications and the expertise and workforce development efforts urgently required to shift integrated systems toward bioenergy research more broadly. Participants discussed many existing and future applications of AI/ML for biosystems design ranging from enzymes to plants and microbes, microbiomes, and bioprocess development. They also identified three key categories of scientific and technical opportunities and challenges: high-quality data, AI/ML algorithms, and laboratory automation. Several main takeaways emerged from the workshop: 1. Numerous AI/ML and automated experimentation applications exist for a variety of DOE mission needs in energy and the environment; 2. Exemplary research grand challenges for which AI/ML could provide solutions include: building microbes and microbial communities to specifications, developing closed-loop autonomous design and control for biosystems design, and advancing scale-up and automation; 3. Lack of sufficient high-quality, annotated data hinders the development of AI/ML applications; 4. New and improved AI/ML tools are needed, particularly those meeting the specific needs of the BER and BETO research communities; 5. Trade-offs in performance, cost, and reliability exist between deploying commercially available versus building custom-developed instrumentation and software for automated or autonomous experimentation; translation of manual to automated or autonomous methods is often a nontrivial endeavor; 6. Training a new generation of young scientists who can develop and apply AI/ML tools is needed to solve long-standing scientific challenges in bioenergy research. The integration of AI/ML tools and automated experimentation represents a new data-driven research paradigm complementary to the traditional hypothesis-driven research paradigm. This paradigm accelerates design and optimization of biological systems and processes for a variety of DOE mission needs in energy and the environment. The AMBER workshop broadly explored the potential of this new paradigm for bioenergy research, of particular interest to BER and BETO, and identified key challenges and opportunities that DOE can address in the coming years by leveraging its unique capabilities and resources.

59 BASIC BIOLOGICAL SCIENCES↗