Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

MINE: maximally informative next experiment—toward a new GWAS experimental design and methodology

Abstract The computational methodology of Genome Wide Association Studies (GWAS) currently has several limitations: (i) the number of observations (rows) on a quantitative trait tends to be smaller than the number of single nucleotide polymorphisms (SNPs) (columns) in the design matrix; (ii) each SNP is usually modeled separately, failing to acknowledge interaction between each other (ie epistasis); (iii) there is implicit linkage disequilibrium (LD) between neighboring SNPs due to their linkage. To overcome these issues, we developed a tool that uses ensemble methods to fit mixed linear models to GWAS data, and these ensemble methods include the development of a new experimental design approach in GWAS, which uses the resultant models and data to select the next informative experiment over time. This new adaptive and staged approach for GWAS experimental design was developed and tested in a 3 yr adaptive model-guided discovery experiment against a fixed classical design. In Sorghum bicolor a total of 79, 86, and 78 accessions were tested in years 1, 2, and 3, respectively out of 343 accessions available in the Bioenergy Association Panel (BAP) each identified for 232,303 SNPs, 1 every 2–3 kb in the genomes. We demonstrated the feasibility of MINE enacted with 8 people in the field per year over 3 yr vs in 1 large classical design enacted with 20 people in 1 yr. The MINE results for chromosomal regions identified controlling dry weight were confirmed against results from previous sorghum GWAS experiments and 1 large classical design for the BAP panel.

Genetics & Heredity↗

AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing

The recent development of deep learning has been mostly focusing on Euclidean data, such as images, videos, audios, etc. However, most real-world information and relation are often expressed as graphs. To efficiently learn from graph data, graph convolutional networks (GCNs) emerge as a promising approach, showing advantages in several practical applications such as social network analysis, knowledge discovery, 3D modeling, motion capturing, etc. Real-world graphs are usually extremely large and imbalanced, posting significant performance demand and design challenges on the hardware dedicated for GCN inference. In this paper, we propose an architecture design called UW-GCN to accelerate graph convolutional network inference. To tackle the major performance bottleneck from workload imbalance, we propose dynamic neighborhood stealing and remote chunk shuffling techniques, relying on hardware flexibility to achieve hardware auto-tuning under negligible area or delay overhead. Specifically, UW-GCN is able to smartly profile the sparse graph pattern while continuously adjusting the workload distribution via routing reconfiguration among parallel processing elements (PEs). The ideal configuration is then reused in the remaining iterations. To the best of our knowledge, this is the first accelerator design particularly for GCN and the first work relying on hardware auto-tuning, which is normally based on software, to achieve near-optimal workload balance in processing sparse structures.

Geng, Tong↗

Root system architecture in cereals: progress, challenges and perspective

We report roots are essential multifunctional plant organs involved in water and nutrient uptake, metabolite storage, anchorage, mechanical support, and interaction with the soil environment. Understanding of this ‘hidden half’ provides potential for manipulation of root system architecture (RSA) traits to optimize resource use efficiency and grain yield in cereal crops. Unfortunately, root traits are highly neglected in breeding due to the challenges of phenotyping, but could have large rewards if the variability in RSA traits can be fully exploited. Until now, a plethora of genes have been characterized in detail for their potential role in improving RSA. The use of forward genetics approaches to find sequence variations in genes underpinning desirable RSA would be highly beneficial. Advances in computer vision applications have allowed image-based approaches for high-throughput phenotyping of RSA traits that can be used by any laboratory worldwide to make progress in understanding root function and dissection of the genetics. At the same time, the frontiers of root measurement include non-invasive methods like X-ray computer tomography and magnetic resonance imaging that facilitate new types of temporal studies. Root physiology and ecology are further supported by spatiotemporal root simulation modeling. The discovery of component traits providing improved resilience and yield advantage in target environments is a key necessity for mainstreaming root-based cereal breeding. The integrated use of pan-genome resources, now available in most cereals, coupled with new in-field phenotyping platforms has the potential for precise selection of superior genotypes with improved RSA.

59 BASIC BIOLOGICAL SCIENCES↗

Transforming ENERGY through Computational Excellence

Computational methods underpin advancing the science and engineering of energy efficiency, sustainable transportation, renewable power technologies, and developing a knowledge base to optimize energy systems. Researchers with access to enough computing, and the right type, can focus their ingenuity and creativity on addressing the energy challenges. NREL’s advanced computing influence spans several common themes across the Office of Energy Efficiency and Renewable Energy (EERE), including materials discovery, process modeling, fluid dynamics, resource mapping, and analysis of large-scale systems with real-time optimization.

advanced computing↗

Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery

Large language models (LLMs) are evolving from chatbots with limited tool-using capabilities to agentic AI systems that can perform deep research, assist in proposing hypotheses, help design experiments, automate data analysis, and draft scientific reports. However, there are currently two bottlenecks limiting LLMs' real-world impact on the broader scientific research community beyond academic demonstrations: lack of interoperability (repetitive manual tool-integration is required across scenarios) and the need for scalable coordination (unstructured communication and memory become brittle as the number of agents grows). In this Perspective, we argue that the next phase of agentic scientific discovery requires the development of an ecosystem of protocol-native agents and tools organized through hierarchies inspired by human society, beyond the current paradigm of a single monolithic “AI scientist”. We use Model Context Protocol (MCP) as a concrete example of an emerging interoperability layer for scientific tool and context exchange, and we propose three complementary pathways to increase the scaling capabilities of an MCP-native scientific ecosystem by addressing the composability issues: (1) MCP servers for high-value scientific tools maintained by domain experts, (2) automated transformation of existing code repositories into MCP services, and (3) autonomous invention and evolution of new agents and workflows. Finally, we provide a practical roadmap for scaling AI-driven scientific discovery by expanding tool supply and coordination in MCP-native scientific ecosystems.

97 MATHEMATICS AND COMPUTING↗

A network-enabled pipeline for gene discovery and validation in non-model plant species

Identifying key regulators of important genes in non-model crop species is challenging due to limited multi-omics resources. To address this, we introduce the network-enabled gene discovery pipeline NEEDLE, a user-friendly tool that systematically generates coexpression gene network modules, measures gene connectivity, and establishes network hierarchy to pinpoint key transcriptional regulators from dynamic transcriptome datasets. After validating its accuracy with two independent datasets, we applied NEEDLE to identify transcription factors (TFs) regulating the expression of cellulose synthase-like F6 ( CSLF6 ), a crucial cell wall biosynthetic gene, in Brachypodium and sorghum. Our analyses uncover regulators of CSLF6 and also shed light on the evolutionary conservation or divergence of gene regulatory elements among grass species. These results highlight NEEDLE’s capability to provide biologically relevant TF predictions and demonstrate its value for non-model plant species with dynamic transcriptome datasets.

59 BASIC BIOLOGICAL SCIENCES↗

A detailed map of Higgs boson interactions by the ATLAS experiment ten years after the discovery

The standard model of particle physics describes the known fundamental particles and forces that make up our Universe, with the exception of gravity. One of the central features of the standard model is a field that permeates all of space and interacts with fundamental particles. The quantum excitation of this field, known as the Higgs field, manifests itself as the Higgs boson, the only fundamental particle with no spin. In 2012, a particle with properties consistent with the Higgs boson of the standard model was observed by the ATLAS and CMS experiments at the Large Hadron Collider at CERN. Since then, more than 30 times as many Higgs bosons have been recorded by the ATLAS experiment, enabling much more precise measurements and new tests of the theory. Here, on the basis of this larger dataset, we combine an unprecedented number of production and decay processes of the Higgs boson to scrutinize its interactions with elementary particles. Interactions with gluons, photons, and W and Z bosons—the carriers of the strong, electromagnetic and weak forces—are studied in detail. Interactions with three third-generation matter particles (bottom (b) and top (t) quarks, and tau leptons (τ)) are well measured and indications of interactions with a second-generation particle (muons, μ) are emerging. These tests reveal that the Higgs boson discovered ten years ago is remarkably consistent with the predictions of the theory and provide stringent constraints on many models of new phenomena beyond the standard model.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

First principles reaction discovery: from the Schrodinger equation to experimental prediction for methane pyrolysis

Our recent success in exploiting graphical processing units (GPUs) to accelerate quantum chemistry computations led to the development of the ab initio nanoreactor, a computational framework for automatic reaction discovery and kinetic model construction. In this work, we apply the ab initio nanoreactor to methane pyrolysis, from automatic reaction discovery to path refinement and kinetic modeling. Elementary reactions occurring during methane pyrolysis are revealed using GPU-accelerated ab initio molecular dynamics simulations. Subsequently, these reaction paths are refined at a higher level of theory with optimized reactant, product, and transition state geometries. Reaction rate coefficients are calculated by transition state theory based on the optimized reaction paths. The discovered reactions lead to a kinetic model with 53 species and 134 reactions, which is validated against experimental data and simulations using literature kinetic models. We highlight the advantage of leveraging local brute force and Monte Carlo sensitivity analysis approaches for efficient identification of important reactions. Both sensitivity approaches can further improve the accuracy of the methane pyrolysis kinetic model. The results in this work demonstrate the power of the ab initio nanoreactor framework for computationally affordable systematic reaction discovery and accurate kinetic modeling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Leveraging the Higgs to Discover Physics Beyond the Standard Model (Final Technical Report)

The discovery of an apparently Standard Model-like Higgs at the Large Hadron Collider (LHC) heralds the start of a new era in particle physics. While the Higgs marks the completion of the Standard Model framework, it offers far more opportunities in the search for physics beyond the Standard Model. The Higgs boson raises a pressing theoretical problem known as the hierarchy problem: why is an elementary scalar particle so light when quantum corrections tie its mass to the highest energy scales? The Higgs also provides an unprecedented experimental opportunity as a bellwether of new physics: it may be merely the first of several states in the electroweak symmetry breaking sector, while its production and decays may provide unique evidence for additional particles. Research supported by this award leveraged the Higgs boson to explore new physics from both directions, developing novel approaches to solving the hierarchy problem posed by the Higgs boson and directly employing the Higgs as a new tool for discovery. Given that null results at the LHC and other experiments have begun to endanger conventional approaches to the hierarchy problem, research supported by the award identified original solutions to the hierarchy problem wherein the lightest degrees of freedom protecting the Higgs boson carry no Standard Model quantum numbers and thus evade existing searches. The PI's approach combined standard tools of quantum field theory with novel applications of the orbifold reduction of continuous symmetries to define the framework of "neutral naturalness'' and explore its experimental consequences across the energy, intensity, and cosmic frontiers. In employing the Higgs directly as a tool for discovery, research supported by this award articulated a systematic approach to searching for extensions of the Higgs sector at the LHC and pursued four key avenues through which the Higgs can be used to uncover new physics across a range of experiments: (1) as a direct final state probe; (2) as an indirect probe through its couplings; (3) as a portal to states neutral under the Standard Model; and (4) as a source of exotic processes in displaced decays.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Machine Learning for a-posteriori model-observed data fusion to enhance predictive value of ESM output

The proposed research is consistent with focal areas 2 and 3: when successful, this effort will enhance the predictive value of Earth System Model (ESM) output by identifying and quantifying its systematic deviations from available observations and by correcting and recalibrating to existing data. The product is a hybrid process & data-driven modeling tool whose estimation can also bring insights through pattern discovery recognition on model deficiencies on the one hand and data gaps on the other.

54 ENVIRONMENTAL SCIENCES↗

Data-driven models of nonautonomous systems

Nonautonomous dynamical systems are characterized by time-dependent inputs, which complicates the discovery of predictive models describing the spatiotemporal evolution of the state variables of quantities of interest from their temporal snapshots. When dynamic mode decomposition (DMD) is used to infer a linear model, this difficulty manifests itself in the need to approximate the time-dependent Koopman operators. Our approach is to approximate the original nonautonomous system with a modified system derived via a local parameterization of the time-dependent inputs. The modified system comprises a sequence of local parametric systems, which are subsequently approximated by a parametric surrogate model using the DRIPS (dimension reduction and interpolation in parameter space) framework. The offline step of DRIPS relies on DMD to build a linear surrogate model, endowed with reduced-order bases for the observables mapped from training data. The online step interpolates on suitable manifolds to construct a sequence of iterative parametric surrogate models; the target/test parameter points on these manifolds are specified by a local parameterization of the test time-dependent inputs. Here, we use numerical experimentation to demonstrate the robustness of our method and compare its performance with that of deep neural networks.

97 MATHEMATICS AND COMPUTING↗

A framework for strategic discovery of credible neural network surrogate models under uncertainty

The widespread integration of deep neural networks in developing data-driven surrogate models for high-fidelity simulations of complex physical systems highlights the critical necessity for robust uncertainty quantification techniques and credibility assessment methodologies, ensuring the reliable deployment of surrogate models in consequential decision-making. Here, this study presents the Occam Plausibility Algorithm for surrogate models (OPAL-surrogate), providing a systematic framework to uncover predictive neural network-based surrogate models within the large space of potential models, including various neural network classes and choices of architecture and hyperparameters. The framework is grounded in hierarchical Bayesian inferences and employs model validation tests to evaluate the credibility and prediction reliability of the surrogate models under uncertainty. Leveraging these principles, OPAL-surrogate introduces a systematic and efficient strategy for balancing the trade-off between model complexity, accuracy, and prediction uncertainty. The effectiveness of OPAL-surrogate is demonstrated through two modeling problems, including the deformation of porous materials for building insulation and turbulent combustion flow for ablation of solid fuels within hybrid rocket motors.

42 ENGINEERING↗

DLHub: Simplifying publication, discovery, and use of machine learning models in science

Machine Learning (ML) has become a critical tool enabling new methods of analysis and driving deeper understanding of phenomena across scientific disciplines. There is a growing need for "learning systems" to support various phases in the ML lifecycle. While others have focused on supporting model development, training, and inference, few have focused on the unique challenges inherent in science, such as the need to publish and share models and to serve them on a range of available computing resources. In this paper, we present the Data and Learning Hub for science (DLHub), a learning system designed to support these use cases. Specifically, DLHub enables publication of models, with descriptive metadata, persistent identifiers, and flexible access control. It packages arbitrary models into portable servable containers, and enables low-latency, distributed serving of these models on heterogeneous compute resources. In this work, we show that DLHub supports low-latency model inference comparable to other model serving systems including TensorFlow Serving, SageMaker, and Clipper, and improved performance, by up to 95%, with batching and memoization enabled. We also show that DLHub can scale to concurrently serve models on 500 containers. Finally, we describe five case studies that highlight the use of DLHub for scientific applications.

97 MATHEMATICS AND COMPUTING↗