Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery

Large language models (LLMs) are evolving from chatbots with limited tool-using capabilities to agentic AI systems that can perform deep research, assist in proposing hypotheses, help design experiments, automate data analysis, and draft scientific reports. However, there are currently two bottlenecks limiting LLMs' real-world impact on the broader scientific research community beyond academic demonstrations: lack of interoperability (repetitive manual tool-integration is required across scenarios) and the need for scalable coordination (unstructured communication and memory become brittle as the number of agents grows). In this Perspective, we argue that the next phase of agentic scientific discovery requires the development of an ecosystem of protocol-native agents and tools organized through hierarchies inspired by human society, beyond the current paradigm of a single monolithic “AI scientist”. We use Model Context Protocol (MCP) as a concrete example of an emerging interoperability layer for scientific tool and context exchange, and we propose three complementary pathways to increase the scaling capabilities of an MCP-native scientific ecosystem by addressing the composability issues: (1) MCP servers for high-value scientific tools maintained by domain experts, (2) automated transformation of existing code repositories into MCP services, and (3) autonomous invention and evolution of new agents and workflows. Finally, we provide a practical roadmap for scaling AI-driven scientific discovery by expanding tool supply and coordination in MCP-native scientific ecosystems.

97 MATHEMATICS AND COMPUTING↗

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES↗

FAST Discovery of a Fast Neutral Hydrogen Outflow

Abstract In this letter, we report the discovery of a fast neutral hydrogen outflow in SDSS J145239.38+062738.0, a merging radio galaxy containing an optical type I active galactic nucleus (AGN). This discovery was made through observations conducted by the Five-hundred-meter Aperture Spherical radio Telescope (FAST) using redshifted 21 cm absorption. The outflow exhibits a blueshifted velocity likely up to ∼−1000 km s −1 with respect to the systemic velocity of the host galaxy with an absorption strength of ∼−0.6 mJy beam −1 corresponding to an optical depth of 0.002 atv= −500 km s −1 . The mass outflow rate ranges between 2.8 × 10 −2 and 3.6M ⊙ yr −1 , implying an energy outflow rate ranging between 4.2 × 10 39 and 9.7 × 10 40 erg s −1 , assuming 100 K s< 1000 K. Plausible drivers of the outflow include the starbursts, AGN radiation, and radio jet, the last of which is considered the most likely culprit according to the kinematics. By analyzing the properties of the outflow, AGN, and jet, we find that if the Hioutflow is driven by the AGN radiation, the AGN radiation does not seem powerful enough to provide negative feedback, whereas the radio jet shows the potential to provide negative feedback. Our observations contribute another example of a fast outflow detected in neutral hydrogen and demonstrate the capability of FAST in detecting such outflows.

Astronomy & Astrophysics↗

Discovery and Spectroscopic Characterization of a Distant, Compact Milky Way Satellite in Gemini

We present the discovery of a compact Milky Way satellite in the constellation of Gemini. This system was discovered by cross-matching detections from two independent search algorithms applied to Blanco/DECam data from the third data release of the DECam Local Volume Exploration survey (DELVE DR3), and confirmed with deeper imaging from Gemini/GMOS-N. Based on these data, we determine that the system is an ultra-faint ($M_V = -2.1^{+0.4}_{-0.6}$), compact ($r_{1/2} = 8.6^{+1.4}_{-1.2}$ pc) system located at a heliocentric distance of $120^{+7}_{-6}$ kpc. These physical properties place the system in the regime of ambiguous, ultra-faint compact Milky Way halo satellites that cannot be confidently classified as dwarf galaxies or star clusters from morphology alone; we therefore name the system DELVE 8/Gemini I. From medium-resolution Keck/DEIMOS spectroscopy, we securely identify four members including two blue horizontal branch stars, confirming the system as a bound satellite moving at a mean radial velocity of $v_{\rm hel} = -82.7^{+3.7}_{-3.9} {\rm km\,s}^{-1}$. We also use these spectra to place an upper limit of $\rm [Fe/H] \lesssim -2.5$ on the metallicity of DELVE 8/Gemini I's brightest star, supporting the classification of the system as either an ancient star cluster or ultra-faint dwarf galaxy. The discovery of faint, distant systems similar to DELVE 8/Gemini I is expected to become more common with upcoming surveys.

Overdeck, K. [Chicago U., Astron. Astrophys. Ctr.;↗

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery↗

A network-enabled pipeline for gene discovery and validation in non-model plant species

Identifying key regulators of important genes in non-model crop species is challenging due to limited multi-omics resources. To address this, we introduce the network-enabled gene discovery pipeline NEEDLE, a user-friendly tool that systematically generates coexpression gene network modules, measures gene connectivity, and establishes network hierarchy to pinpoint key transcriptional regulators from dynamic transcriptome datasets. After validating its accuracy with two independent datasets, we applied NEEDLE to identify transcription factors (TFs) regulating the expression of cellulose synthase-like F6 ( CSLF6 ), a crucial cell wall biosynthetic gene, in Brachypodium and sorghum. Our analyses uncover regulators of CSLF6 and also shed light on the evolutionary conservation or divergence of gene regulatory elements among grass species. These results highlight NEEDLE’s capability to provide biologically relevant TF predictions and demonstrate its value for non-model plant species with dynamic transcriptome datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Language models for materials discovery and sustainability: Progress, challenges, and opportunities

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI’s GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. Here, in this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.

36 MATERIALS SCIENCE↗

Targeted materials discovery using Bayesian algorithm execution

Rapid discovery and synthesis of future materials requires intelligent data acquisition strategies to navigate large design spaces. A popular strategy is Bayesian optimization, which aims to find candidates that maximize material properties; however, materials design often requires finding specific subsets of the design space which meet more complex or specialized goals. We present a framework that captures experimental goals through straightforward user-defined filtering algorithms. These algorithms are automatically translated into one of three intelligent, parameter-free, sequential data collection strategies (SwitchBAX, InfoBAX, and MeanBAX), bypassing the time-consuming and difficult process of task-specific acquisition function design. Our framework is tailored for typical discrete search spaces involving multiple measured physical properties and short time-horizon decision making. We demonstrate this approach on datasets for TiO 2 nanoparticle synthesis and magnetic materials characterization, and show that our methods are significantly more efficient than state-of-the-art approaches. Overall, our framework provides a practical solution for navigating the complexities of materials design, and helps lay groundwork for the accelerated development of advanced materials.

42 ENGINEERING↗

Discovery of hybrid chemical synthesis pathways with DORAnet

Developing efficient tools for discovering novel synthesis pathways is essential to advance chemical production methods that maximize the use of resources and energy. We introduce DORAnet (Designing Optimal Reaction Avenues Network Enumeration Tool), an open-source computational framework that addresses key limitations in current computer-aided synthesis planning (CASP) tools. DORAnet integrates both chemical/chemocatalytic (i.e., non-enzymatic) and enzymatic transformations, enabling the discovery of hybrid synthesis pathways. With 390 expert-curated chemical/chemocatalytic reaction rules and 3606 enzymatic rules derived from MetaCyc, it provides extensive flexibility for synthetic chemists and biotechnologists. The framework features customizable network expansion strategies, advanced filtering, and pathway search, ranking, and visualization tools. Validated against known reaction data, DORAnet successfully identified both established and novel synthesis routes for key industrial chemicals. In a case study involving 51 high-volume targets, DORAnet frequently ranked known commercial pathways among the top three results, demonstrating its practical relevance and ranking accuracy, while also uncovering numerous alternative (hybrid) synthesis pathways that were highly ranked.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tailoring Molecular Space to Navigate Phase Complexity in Cs-Based Quasi-2D Perovskites via Gated-Gaussian-Driven High-Throughput Discovery

Cesium-based quasi-2D halide perovskites (HPs) offer promising functionalities and low-temperature manufacturability, suited to stable tandem photovoltaics. However, the chemical interplays between the molecular spacers and the inorganic building blocks during crystallization cause substantial phase complexities in the resulting matrices. To successfully optimize and implement the quasi-2D HP functionalities, a systematic understanding of spacer chemistry, along with the seamless navigation of the inherently discrete molecular space, is necessary. Herein, by utilizing high-throughput automated experimentation, the phase complexities in the molecular space of quasi-2D HPs are explored, thus identifying the chemical roles of the spacer cations on the synthesis and functionalities of the complex materials. Furthermore, a novel active machine learning algorithm leveraging a two-stage decision-making process, called gated Gaussian process Bayesian optimization is introduced, to navigate the discrete ternary chemical space defined with two distinctive spacer molecules. Through simultaneous optimization of photoluminescence intensity and stability that “tailors” the chemistry in the molecular space, a ternary-compositional quasi-2D HP film realizing excellent optoelectronic functionalities is demonstrated. Finally, this work not only provides a pathway for the rational and bespoke design of complex HP materials but also sets the stage for accelerated materials discovery in other multifunctional systems.

36 MATERIALS SCIENCE↗

Biocatalyst discovery and design for plastics deconstruction: A multi‐scale perspective

Plastic waste accumulation poses significant environmental challenges due to a lack of economical solutions for the molecular deconstruction of diverse synthetic polymers. Biological‐based degradation offers promise but is hindered by the crystallinity, hydrophobicity, and additive complexity of plastics, which restrict biocatalyst access and activity. To address these problems, we propose a multi‐scale framework that combines detailed materials characterization, optimization of plastic‐biomolecular interfacial interactions, and enhancement of biocatalytic kinetics to develop effective plastic‐deconstructing enzymes. This approach leverages principles from reaction kinetics, transport and interfacial phenomena, and enzyme engineering to systematically address barriers across diverse plastic types. Our framework aims to accelerate the discovery and optimization of biocatalysts capable of scalable, selective, and efficient deconstruction of plastic waste. These advances hold potential to enable sustainable biological recycling and upcycling pathways, contributing to global efforts in mitigating plastic pollution and promoting circular material economies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Targeted Chemical Looping Materials Discovery by an Inverse Design

Chemical looping with oxygen uncoupling (CLOU) materials is actively sought for combustion of carbonaceous materials to achieve complete conversion and capture of carbon dioxide. These materials may play a vital role in reducing atmospheric carbon via negative carbon output. However, there is no one‐size‐fits‐all approach as different operating conditions and feedstocks may require different CLOU materials. As a result, the exploration and discovery of high‐performance CLOU materials can be a slow process. To address this challenge, a high‐throughput inverse machine learning workflow that identifies optimum materials from perovskite oxides for a given set of targets is developed—temperature and Gibbs free energy of oxygen formation. The model is trained on high‐throughput density functional theory calculations of CLOU materials and inverts the materials design process using a genetic algorithm to produce realistic substituted SrFeO 3‐δ compositions as output. Using the inverse model, it is able to identify several interesting new families of CLOU materials: Sr 1‐ x A x Fe 1‐ y B y O 3‐δ (e.g., A = Ca or K; B = Mg, Bi, Mn, Ni, Co, Cu, or Zn). These materials have shown promising properties, and some of them even outperform the benchmark material in terms of oxygen release kinetics under relevant CLOU operating conditions.

36 MATERIALS SCIENCE↗

Machine Learning‐Guided Discovery of High‐Entropy Perovskite Oxide Electrocatalysts via Oxygen Vacancy Engineering

Abstract High‐entropy perovskite oxides (HEPOs) have recently emerged as multifunctional catalysts. However, the HEPOs’ structural and compositional complexity hinders the easy and accurate extrapolation of activity indicators, which are essential for establishing structure‐property correlations. Here, OxiGraphX, is introduced as a novel graph neural network (GNN) model designed to capture the complex relationships among structure, composition, and atomic chemical environments for accurate prediction of oxygen vacancy formation energies (OVFEs) in HEPOs. By integrating machine learning (ML), density functional theory (DFT), and experimental validation, this work demonstrates an efficient framework for rapidly and accurately screening HEPO electrocatalysts for oxygen evolution reaction (OER). The OxiGraphX predicts OVFEs with a precision exceeding existing data, enabling the identification of compositions of higher oxygen vacancy content (OVC) and, thus, higher catalytic activity. Furthermore, the model explores latent spaces that translate effectively into experimental domains, bridging computational predictions with real‐world applications. This approach accelerates the discovery of high‐performance HEPO catalysts while providing deeper insights into their catalytic mechanisms.

Chemistry↗

Discovery of nuclear isomers

Currently, 1917 nuclear isomers with half-lives longer than 100 ns have been discovered in 1310 different nuclides and 103 different elements. Here, while the physical properties of isomers have been compiled before, this is the first compilation of the isomer discoveries. For each isomer the reference, year, laboratory and country are documented.

Thoennessen, M. [Michigan State University, East L↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗

Reimagining metal-organic framework discovery: Integrating experiment, computation, and artificial intelligence

The traditional development of novel metal–organic frameworks (MOFs) is often hindered by challenges such as synthetic accessibility and time- and resource-intensive experimentation. High-throughput, automated experimental and computational techniques have enabled rapid chemical space exploration and theoretical MOF design. When combined with artificial intelligence (AI), these methods can be used to lead autonomous laboratories to new frontiers for MOF discovery, where these materials can be designed for a specific application, efficiently synthesized, characterized, and evaluated. Here, this perspective highlights the role of AI in advancing automated MOF synthesis and characterization, computational MOF design and screening, and the integration of these approaches within autonomous workflows to ultimately enable the MOF laboratories of the future.

Gaidimas, Madeleine A. [Northwestern University, E↗

Machine learning enabled discovery of superhard and ultrahard carbon polymorphs

The demand for multifunctional materials has motivated the move from near-equilibrium materials to metastable i.e. out-of-equilibrium phases that can meet several desired target properties. The search for such metastable phases with exotic properties is non-trivial and often serendipitous. Inverse design approaches based on evolutionary search have been powerful tools, but such traditional searches have focused on identifying primarily stable and metastable materials with the lowest enthalpy. The inverse design of materials, with a focus on a desired property such as, for example, hardness is a challenging task because of the expensive computational cost involved in sampling multiple structures. The recent advances in machine learning have brought new powerful AI techniques to the forefront which can potentially revolutionize the inverse design and discovery of materials, especially metastable phases capable of meeting multifunctionality. Here, in this work, we develop and apply an automated reinforcement learning workflow for inverse design that integrates first principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the superhard and ultrahard metastable phases of Carbon. We demonstrate an automatic machine learning based inverse design workflow to map new undiscovered metastable states ranging from near equilibrium to those far-from-equilibrium that satisfy multiple property objectives, specifically bulk moduli, shear moduli and hardness. We create a comprehensive library of carbon stable and metastable phases with varying hardness and subsequently shortlist 10 top performing candidate carbon structures, including two newly reported phases, based on their hardness and characterize their temperature dependent mechanical properties. A neural network model is built using featurization of allotropes of carbon to predict the quasi-harmonic Gibbs free energies. The Gibbs free energies of the top performing phases are analyzed to get an estimate of the experimental synthesizability of these superhard and ultrahard carbon phases. In general, we show using machine learning based inverse design approaches how hitherto inaccessible metastable states can be identified and potentially synthesized to meet the demand for multifunctional materials.

Balasubramanian, Karthik [Univ. of Illinois, Chica↗

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution↗