Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Health Management Applications for International Space Station

Traditional mission and vehicle management involves teams of highly trained specialists monitoring vehicle status and crew activities, responding rapidly to any anomalies encountered during operations. These teams work from the Mission Control Center and have access to engineering support teams with specialized expertise in International Space Station (ISS) subsystems. Integrated System Health Management (ISHM) applications can significantly augment these capabilities by providing enhanced monitoring, prognostic and diagnostic tools for critical decision support and mission management. The Intelligent Systems Division of NASA Ames Research Center is developing many prototype applications using model-based reasoning, data mining and simulation, working with Mission Control through the ISHM Testbed and Prototypes Project. This paper will briefly describe information technology that supports current mission management practice, and will extend this to a vision for future mission control workflow incorporating new ISHM applications. It will describe ISHM applications currently under development at NASA and will define technical approaches for implementing our vision of future human exploration mission management incorporating artificial intelligence and distributed web service architectures using specific examples. Several prototypes are under development, each highlighting a different computational approach. The ISStrider application allows in-depth analysis of Caution and Warning (C&W) events by correlating real-time telemetry with the logical fault trees used to define off-nominal events. The application uses live telemetry data and the Livingstone diagnostic inference engine to display the specific parameters and fault trees that generated the C&W event, allowing a flight controller to identify the root cause of the event from thousands of possibilities by simply navigating animated fault tree models on their workstation. SimStation models the functional power flow for the ISS Electrical Power System and can predict power balance for nominal and off-nominal conditions. SimStation uses realtime telemetry data to keep detailed computational physics models synchronized with actual ISS power system state. In the event of failure, the application can then rapidly diagnose root cause, predict future resource levels and even correlate technical documents relevant to the specific failure. These advanced computational models will allow better insight and more precise control of ISS subsystems, increasing safety margins by speeding up anomaly resolution and reducing,engineering team effort and cost. This technology will make operating ISS more efficient and is directly applicable to next-generation exploration missions and Crew Exploration Vehicles.

Alena, Richard↗

Impact of the International Space Station Research Results

The International Space Station (ISS) facilitates research that benefits human lives on Earth and serves as the primary testing ground for technology development to sustain life in the extreme environment of space. To date, investigators have published a wide range of ISS science results, from improved theories about the creation of stars to the outcome of data mining “omics” repositories of previously completed ISS investigations. Because of the unique microgravity environment of the ISS laboratory and the multidisciplinary and international nature of the research, analyzing ISS scientific impacts is an exceptional challenge. As a result, the ISS Program Science Forum (PSF), made up of senior science representatives across the ISS international partnership, uses various methods to describe the impacts of ISS research activities. For the most part, past papers written by PSF members to assess the overall ISS research impact have focused on exhibiting ISS research impact by quantifying ISS research output or its perceived benefits for humanity. This paper proposes a new assessment of ISS impact from the perspective of the end users’ needs. To that end, the authors use visualizations and metrics of scientific publication data to show the ISS research influence on traditional scientific fields, its global reach and the benefits to people across the globe.

Diallo, Ousmane N.↗

Large language models for batteries

Large Language Models (LLMs) are advanced artificial intelligence systems capable of solving diverse tasks using language, reasoning, and external tools. Despite their growing deployment in academia and industry, their potential remains underexplored in battery research. This review presents a comprehensive overview of existing and emerging applications of LLMs in batterie field, addressing two critical questions: What can LLMs offer to support battery-related tasks, and how to develop more effective models for this purpose. We begin by outlining the principles of LLMs and criteria for selecting appropriate models and tools for battery research and development. We then explore their roles in text-mining, data interpretation, and the development of intelligent battery systems. In parallel, we discuss technical challenges, such as data standardizing and sharing, model evaluation, and tool integration. Lastly, we propose future research directions with short-, medium-, and long-term goals and highlight more broad perspectives for connecting experts and cross-disciplinary collaborations.

SoC↗

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Significant DBSCAN+: Statistically Robust Density-based Clustering

Cluster detection is important and widely used in a variety of applications, including public health, public safety, transportation, and so on. Given a collection of data points, we aim to detect density-connected spatial clusters with varying geometric shapes and densities, under the constraint that the clusters are statistically significant. The problem is challenging, because many societal applications and domain science studies have low tolerance for spurious results, and clusters may have arbitrary shapes and varying densities. As a classical topic in data mining and learning, a myriad of techniques have been developed to detect clusters with both varying shapes and densities (e.g., density-based, hierarchical, spectral, or deep clustering methods). However, the vast majority of these techniques do not consider statistical rigor and are susceptible to detecting spurious clusters formed as a result of natural randomness. On the other hand, scan statistic approaches explicitly control the rate of spurious results, but they typically assume a single “hotspot” of over-density and many rely on further assumptions such as a tessellated input space. To unite the strengths of both lines of work, we propose a statistically robust formulation of a multi-scale DBSCAN, namely Significant DBSCAN+, to identify significant clusters that are density connected. As we will show, incorporation of statistical rigor is a powerful mechanism that allows the new Significant DBSCAN+ to outperform state-of-the-art clustering techniques in various scenarios. We also propose computational enhancements to speed-up the proposed approach. Experiment results show that Significant DBSCAN+ can simultaneously improve the success rate of true cluster detection (e.g., 10–20% increases in absolute F1 scores) and substantially reduce the rate of spurious results (e.g., from thousands/hundreds of spurious detections to none or just a few across 100 datasets), and the acceleration methods can improve the efficiency for both clustered and non-clustered data.

Computer Science↗

Survey on stochastic distribution systems: A full probability density function control theory with potential applications

Complex systems seen either in general engineering practice or economics are subjected to ever increased uncertainties that are mostly represented as random variables or parameters, and the characteristics of random variables are represented by their probability density functions (PDFs). Controlling their PDFs means to shape their stochastic distributions and in general it would provide a full treatment for system analysis and operational control and optimization. This leads to the development of stochastic distribution control (SDC) systems theory in the past decades, where the original aim of the controller design is to realize a shape control of the distributions of certain random variables in their PDFs sense for some engineering processes. Indeed, once the PDFs of these random variables or parameters are used to describe their distribution characters, the control task is to obtain control signals so that the output PDFs of stochastic systems are made to follow their target PDFs. The subject of SDC was initially originated for non-Gaussian stochastic control systems design but has found a wide spectrum of applications in general systems in terms of data-driven modeling, analysis, signal processing (filtering), data mining via multivariable statistics, decision-making (optimization) for systems subjected to uncertainties and even in economics. In this context, SDC constitutes an effective primer tool for complex system analysis, control and operational optimizations. In this review paper, a detailed survey of the developments on the research of SDC systems will be made together with their wide spectrum applications and future perspectives.

42 ENGINEERING↗

4D-STEM Mapping of Nanocrystal Reaction Dynamics and Heterogeneity in a Graphene Liquid Cell

Chemical reaction kinetics at the nanoscale are intertwined with heterogeneity in structure and composition. However, mapping such heterogeneity in a liquid environment is extremely challenging. Here, in this work, we integrate graphene liquid cell (GLC) transmission electron microscopy and four-dimensional scanning transmission electron microscopy to image the etching dynamics of gold nanorods in the reaction media. Critical to our experiment is the small liquid thickness in a GLC that allows the collection of high-quality electron diffraction patterns at low dose conditions. Machine learning-based data-mining of the diffraction patterns maps the three-dimensional nanocrystal orientation, groups spatial domains of various species in the GLC, and identifies newly generated nanocrystallites during reaction, offering a comprehensive understanding on the reaction mechanism inside a nanoenvironment. This work opens opportunities in probing the interplay of structural properties such as phase and strain with solution-phase reaction dynamics, which is important for applications in catalysis, energy storage, and self-assembly.

four-dimensional scanning transmission electron mi↗

Autonomous Synthesis and Inverse Design of Electrochromic Polymers with High Efficiency and Accuracy

Here, the design and synthesis of functional polymers, aimed at targeted properties through specific structures, have long been challenged by their complex and often nonlinear structure–property relationships. Key processes, including knowledge accumulation for predictive design and experimental refinement and validation, are traditionally labor-insensitive and time-consuming, making it difficult to balance accuracy and efficiency. Here, we introduce an accelerated, autonomous system for the on-demand synthesis of electronic polymers that achieves the desired electrochromic functionality with high accuracy and efficiency. Our approach leverages large language model-assisted data mining, a physics-informed copolymer machine learning model, and an AI-driven autonomous robotic workflow in the Polybot lab. Within 72 h, Polybot autonomously synthesized electrochromic polymers (ECPs) with targeted, previously-unreported color values, including green polymers with specific absorption profiles, precisely fine-tuning copolymer structures with a 5% step size in comonomer composition within a three-monomer system. A publicly accessible ECP informatics database has also been created to foster knowledge exchange.

AI-driven Robotic Lab↗

An integrated multi-omics approach identifies the landscape of interferon-α-mediated responses of human pancreatic beta cells

Proinflammatory cytokines are important mediators of pancreatic beta cell dysfunction and demise in the early stages of type 1 diabetes (T1D). Interferon-a (IFNa), a type I interferon member, is expressed in the islets of T1D individuals and it is expression and signaling is regulated by both genetic (T1D risk variants) and environmental factors (viral infections) associated to T1D. We presently characterized human beta cells responses to IFNa by combining ATAC-seq, RNA-seq and proteomics assays. The initial beta cell response to IFNa was characterized by major chromatin remodeling, followed by marked changes in transcriptional and translational regulation. IFNa-induced changes in alternative splicing (AS) and first exon usage increased the diversity of transcripts expressed by beta cells. This, combined with changes observed on protein modification/degradation, ER stress and MHC class I, may significantly expand the peptide repertoire presented by beta cells to the immune system. On the other hand, beta cells up-regulated checkpoint proteins, such as PDL1 and HLA-E, that may protect them against the autoimmune assault. Data mining of the present multi-omics analysis led to the identification of two compound classes that revert IFNa effects on human beta cells and may be translated to clinical trials.

60 APPLIED LIFE SCIENCES↗

Machine learning the metastable phase diagram of covalently bonded carbon

Abstract Conventional phase diagram generation involves experimentation to provide an initial estimate of the set of thermodynamically accessible phases and their boundaries, followed by use of phenomenological models to interpolate between the available experimental data points and extrapolate to experimentally inaccessible regions. Such an approach, combined with high throughput first-principles calculations and data-mining techniques, has led to exhaustive thermodynamic databases (e.g. compatible with the CALPHAD method), albeit focused on the reduced set of phases observed at distinct thermodynamic equilibria. In contrast, materials during their synthesis, operation, or processing, may not reach their thermodynamic equilibrium state but, instead, remain trapped in a local (metastable) free energy minimum, which may exhibit desirable properties. Here, we introduce an automated workflow that integrates first-principles physics and atomistic simulations with machine learning (ML), and high-performance computing to allow rapid exploration of the metastable phases to construct “metastable” phase diagrams for materials far-from-equilibrium. Using carbon as a prototypical system, we demonstrate automated metastable phase diagram construction to map hundreds of metastable states ranging from near equilibrium to far-from-equilibrium (400 meV/atom). We incorporate the free energy calculations into a neural-network-based learning of the equations of state that allows for efficient construction of metastable phase diagrams. We use the metastable phase diagram and identify domains of relative stability and synthesizability of metastable materials. High temperature high pressure experiments using a diamond anvil cell on graphite sample coupled with high-resolution transmission electron microscopy (HRTEM) confirm our metastable phase predictions. In particular, we identify the previously ambiguous structure of n -diamond as a cubic-analog of diaphite-like lonsdaelite phase.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-driven electron-diffraction approach reveals local short-range ordering in CrCoNi with ordering effects

Abstract The exceptional mechanical strength of medium/high-entropy alloys has been attributed to hardening in random solid solutions. Here, we evidence non-random chemical mixing in a CrCoNi alloy, resulting from short-range ordering. A data-mining approach of electron nanodiffraction enabled the study, which is assisted by neutron scattering, atom probe tomography, and diffraction simulation using first-principles theory models. Two samples, one homogenized and one heat-treated, are observed. In both samples, results reveal two types of short-range-order inside nanoclusters that minimize the Cr–Cr nearest neighbors (L1 2 ) or segregate Cr on alternating close-packed planes (L1 1 ). The L1 1 is predominant in the homogenized sample, while the L1 2 formation is promoted by heat-treatment, with the latter being accompanied by a dramatic change in dislocation-slip behavior. These findings uncover short-range order and the resulted chemical heterogeneities behind the mechanical strength in CrCoNi, providing general opportunities for atomistic-structure study in concentrated alloys for the design of strong and ductile materials.

36 MATERIALS SCIENCE↗

Data-driven computational prediction and experimental realization of exotic perovskite-related polar magnets

Rational design of technologically important exotic perovskites is hampered by the insufficient geometrical descriptors and costly and extremely high-pressure synthesis, while the big-data driven compositional identification and precise prediction entangles full understanding of the possible polymorphs and complicated multidimensional calculations of the chemical and thermodynamic parameter space. Here we present a rapid systematic data-mining-driven approach to design exotic perovskites in a high-throughput and discovery speed of the A 2 BB ’O 6 family as exemplified in A 3 TeO 6 . The magnetoelectric polar magnet Co 3 TeO 6 , which is theoretically recognized and experimentally realized at 5 GPa from the six possible polymorphs, undergoes two magnetic transitions at 24 and 58 K and exhibits helical spin structure accompanied by magnetoelastic and magnetoelectric coupling. We expect the applied approach will accelerate the systematic and rapid discovery of new exotic perovskites in a high-throughput manner and can be extended to arbitrary applications in other families.

36 MATERIALS SCIENCE↗

A database of synthetic inelastic neutron scattering spectra from molecules and crystals

Abstract Inelastic neutron scattering (INS) is a powerful tool to study the vibrational dynamics in a material. The analysis and interpretation of the INS spectra, however, are often nontrivial. Unlike diffraction, for which one can quickly calculate the scattering pattern from the structure, the calculation of INS spectra from the structure involves multiple steps requiring significant experience and computational resources. To overcome this barrier, a database of INS spectra consisting of commonly seen materials will be a valuable reference, and it will also lay the foundation of advanced data-driven analysis and interpretation of INS spectra. Here we report such a database compiled for over 20,000 organic molecules and over 10,000 inorganic crystals. The INS spectra are obtained from a streamlined workflow, and the synthetic INS spectra are also verified by available experimental data. The database is expected to greatly facilitate INS data analysis, and it can also enable the utilization of advanced analytics such as data mining and machine learning. Notice: This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a non-exclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan ( http://energy.gov/downloads/doe-public-access-plan ).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Informing selection of drugs for COVID-19 treatment through adverse events analysis

Coronavirus disease 2019 (COVID-19) is an ongoing pandemic and there is an urgent need for safe and effective drugs for COVID-19 treatment. Since developing a new drug is time consuming, many approved or investigational drugs have been repurposed for COVID-19 treatment in clinical trials. Therefore, selection of safe drugs for COVID-19 patients is vital for combating this pandemic. Our goal was to evaluate the safety concerns of drugs by analyzing adverse events reported in post-market surveillance. We collected 296 drugs that have been evaluated in clinical trials for COVID-19 and identified 28,597,464 associated adverse events at the system organ classes (SOCs) level in the FDA adverse events report systems (FAERS). We calculated Z-scores of SOCs that statistically quantify the relative frequency of adverse events of drugs in FAERS to quantitatively measure safety concerns for the drugs. Analyzing the Z-scores revealed that these drugs are associated with different significantly frequent adverse events. Our results suggest that this safety concern metric may serve as a tool to inform selection of drugs with favorable safety profiles for COVID-19 patients in clinical practices. Caution is advised when administering drugs with high Z-scores to patients who are vulnerable to associated adverse events.

59 BASIC BIOLOGICAL SCIENCES↗

Structure determination of the HgcAB complex using metagenome sequence data: insights into microbial mercury methylation

Bacteria and archaea possessing the hgcAB gene pair methylate inorganic mercury (Hg) to form highly toxic methylmercury. HgcA consists of a corrinoid binding domain and a transmembrane domain, and HgcB is a dicluster ferredoxin. However, their detailed structure and function have not been thoroughly characterized. We modeled the HgcAB complex by combining metagenome sequence data mining, coevolution analysis, and Rosetta structure calculations. In addition, we overexpressed HgcA and HgcB in Escherichia coli, confirmed spectroscopically that they bind cobalamin and [4Fe-4S] clusters, respectively, and incorporated these cofactors into the structural model. Surprisingly, the two domains of HgcA do not interact with each other, but HgcB forms extensive contacts with both domains. The model suggests that conserved cysteines in HgcB are involved in shuttling HgII, methylmercury, or both. These findings refine our understanding of the mechanism of Hg methylation and expand the known repertoire of corrinoid methyltransferases in nature.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of processing conditions on the film formation of lead-free halide double perovskite Cs 2 AgBiBr 6

Lead-free halide double perovskites with enhanced stability have gained attention as a promising environmentally friendly alternative to lead-based halide perovskites. Amongst different halide double perovskites, Cs 2 AgBiBr 6 has shown attractive optoelectronic properties and stability, making it a promising candidate for stable high-efficiency optoelectronic devices. Motivated by a data mining effort, we present here in this study the effects of different processing strategies on the microstructure and thin-film formation dynamics of Cs 2 AgBiBr 6 . We apply some of the most successfully used solvent engineering approaches from the halide perovskite research to halide double perovskites, namely antisolvent- and additive-assisted synthesis. Using in situ spectroscopy and diffraction, the film formation of Cs 2 AgBiBr 6 is investigated during spin coating, and the subsequent post-deposition thermal annealing. Dropping antisolvents during spin coating induces immediate supersaturation and crystallization of the wet film, whereas the time of dropping the antisolvent has implications on the film formation dynamics and the final microstructure. For additive (HBr)-assisted synthesis, we show how the addition of HBr affects colloid formation in solution and thus influences the crystallization pathway during thin-film processing. Finally, HBr additive simplifies synthesis in that it doesn't require solution and substrate preheating to obtain pinhole-free films even with higher thickness.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bioinformatic and experimental evidence for suicidal and catalytic plant THI4s

Like fungi and some prokaryotes, plants use a thiazole synthase (THI4) to make the thiazole precursor of thiamin. Fungal THI4s are suicide enzymes that destroy an essential active-site Cys residue to obtain the sulfur atom needed for thiazole formation. In contrast, certain prokaryotic THI4s have no active-site Cys, use sulfide as sulfur donor, and are truly catalytic. The presence of a conserved active-site Cys in plant THI4s and other indirect evidence implies that they are suicidal. To confirm this, we complemented the Arabidopsistz-1 mutant, which lacks THI4 activity, with a His-tagged Arabidopsis THI4 construct. LC–MS analysis of tryptic peptides of the THI4 extracted from leaves showed that the active-site Cys was predominantly in desulfurated form, consistent with THI4 having a suicide mechanism in planta. Unexpectedly, transcriptome data mining and deep proteome profiling showed that barley, wheat, and oat have both a widely expressed canonical THI4 with an active-site Cys, and a THI4-like paralog (non-Cys THI4) that has no active-site Cys and is the major type of THI4 in developing grains. Transcriptomic evidence also indicated that barley, wheat, and oat grains synthesize thiamin de novo, implying that their non-Cys THI4s synthesize thiazole. Structure modeling supported this inference, as did demonstration that non-Cys THI4s have significant capacity to complement thiazole auxotrophy in Escherichia coli. There is thus a prima facie case that non-Cys cereal THI4s, like their prokaryotic counterparts, are catalytic thiazole synthases. Bioenergetic calculations show that, relative to suicide THI4s, such enzymes could save substantial energy during the grain-filling period.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗