Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Structures of glasses created by multiple kinetic arrests

X-ray scattering has been used to characterize glassy itraconazole (ITZ) prepared by cooling at different rates. Faster cooling produces ITZ glasses with lower (or zero) smectic order with more sinusoidal density modulation, larger molecular spacing, and shorter lateral correlation between the rod-like molecules. We find that each glass is characterized by not one, but two fictive temperatures T f (the temperature at which a chosen order parameter is frozen in the equilibrium liquid). The higher T f is associated with the regularity of smectic layers and lateral packing, while the lower T f with the molecular spacings between and within smectic layers. This indicates that different structural features are frozen on different timescales. The two timescales for ITZ correspond to its two relaxation modes observed by dielectric spectroscopy: the slower δ mode (end-over-end rotation) is associated with the freezing of the regularity of molecular packing and the faster α mode (rotation about the long axis) with the freezing of the spacing between molecules. Finally, our finding suggests a way to selectively control the structural features of glasses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GCS site selection in saline Miocene formations in South Louisiana

One of the key requirements of a commercial scale geologic carbon sequestration (GCS) site selection process is identifying injection sites to safely inject and retain injected CO 2 over extended time scales. This requires thorough evaluation of site-specific surface and subsurface features. A site screening and selection framework was developed and used to identify potential sites in South Louisiana, incorporating surface data including site proximity to large CO 2 stationary sources, vicinity of large population centres and state lands, land usage, legacy oil and gas wells, and local/regional geologic trends. Regionally the Miocene sediments thicken and dip towards the south, reflecting a prograding deltaic environment during the Miocene which resulted in the deposition of thick packages of fine-grained sandstones suitable for CO 2 storage. The results of three example selected sites that are spatially separated with varying subsurface and surface features are presented in this study. The multiple stacked sand zones at these sites provide ample amount of pore volume for CO 2 storage with varying degree of stratigraphic/structural traps for long terms storage containment. Here, the screening framework presented in this study and the detailed evaluation of three specific sites in South Louisiana provides a process to screen and characterize sites for GCS projects and the suitability of these Miocene sediments for commercial scale CO 2 storage projects. It is envisioned that the analysis described in this paper will be beneficial for a large number of audiences involved in GCS work and aid in the site selection process.

58 GEOSCIENCES↗

WELLS Interactive Application

The Wellbore Exploration and Location Logistic System (WELLS) Interactive Application is an interactive tool to enable easy exploration and visualization of the living national wellbore database (WELLS Database (https://edx.netl.doe.gov/dataset/wells_database)). The tool and underlying database were created and are maintained by the National Energy Technology Laboratory (NETL), providing visualization of the more than six million public wellbore records from more than 65 authoritative state, federal, and tribal resources. The WELLS Interactive Application serves up wellbore data from oil, gas, underground injection, research, geothermal, geotechnical, groundwater, and other types of wells in a single, standardized, unified system. In addition to the surface location of these wells, the underlying database combines select key attributes for features such as well age, depth, and operating status. The system also provides users with references back to the original sources used in this unified platform. The underlying data can be accessed through the WELLS Database: https://edx.netl.doe.gov/dataset/wells_database Additional Information: The WELLS Interactive Application (formerly titled CO2-Locate) enables visualization and access to the public wellbore records through an intuitive web-based mapping tool. The WELLS Interactive Application was designed to help users visualize, query, analyze, and download wellbore records. Public wellbore points are included as a layer in the Map page, called Public Wells. Additionally, a multivariate hexagon grid summarizing well density from proprietary well data, called Well Density, is included to identify data gaps between the public and proprietary well data. Filtering functionalities in the tool allow these two layers to be spatially filtered by state, county, or basin as well as by status, type, true vertical depth, and spud year. The WELLS Interactive Application also contains a Near Me tool can be used to search and explore wellbore data within a user-defined distance of a specified location on the map, which can also be downloaded. The Query tool allows users to query the selected or filtered wells in the Public Wells layer and export the data. For additional information on these tool functionalities, see the help documentation on the About page of the tool. Notes for Consideration: The Well Density layer provided in this application is derived from proprietary wellbore data, the records of which do not always contain values for key features (status, type, true vertical depth, or spud year). Therefore, data might not be available when layers are queried for all filter combinations. Additionally, visualizing layers and applying filters may take additional time to load (i.e., draw on the map) due to the large size of the data.

ccs↗

Computational Framework for Machine-Learning-Enabled 13 C Fluxomics

13 C metabolic flux analysis (MFA) has emerged as a powerful tool for synthetic biology. This optimization-based approach suffers long computation time and unstable solutions depending on the initial guess. Here, we develop a machine-learning-based framework for 13 C fluxomics. Specifically, training and test data sets are generated by metabolic network decomposition and flux sampling, in which flux ratios at metabolic nodes and simulated labeling patterns of metabolites are used as training targets and features, respectively. To improve prediction accuracy and simplify the model, automated processes are developed for flux ratio selection based on solvability and feature screening based on importance. We found that predictive performance can be significantly improved using both amino acids and central carbon metabolites in comparison with amino acids alone. Together with measured external fluxes, the predicted flux ratios determine the mass balance system, yielding global flux distributions. This approach is validated by flux estimation using both simulated and experimental data in comparison with canonical 13 C MFA. The approach represents a reliable fluxomics method readily applicable to high-throughput metabolic phenotyping, which highlights the advances of intelligent learning algorithms in synthetic biology, specifically in the Test and Learn stage of the Design-Build-Test-Learn cycle.

13C metabolic flux analysis↗

Computationally Guided and Experimentally Validated Design of Custom Chelators for Critical Mineral Recovery

Selective, high throughput separation of target critical metals from complex environments such as fly ash leachates and mining process streams presents a significant challenge for economical production. Custom chelators and sorbents are an attractive technology for selective metal extraction, however it can be difficult to predict their performance, and significant experimental efforts are often required to develop chelating technologies. Here, we present a computational strategy focused on modelling chelator-metal binding interactions and benchmark these results versus experimental data. A computational pipeline combining forcefield, semiempirical, and meta-GGA methods with a thermodynamic framework optimized for error cancellation has been developed to predict binding energies of chelator complexes towards critical mineral recovery applications. This approach, originally validated on [2.2.2] cryptates binding mono- and divalent cations, demonstrated robust predictive capabilities with an R2 of 0.850 against experimental aqueous binding energies. The workflow includes metadynamics for exploring high-dimensional potential energy surfaces and a cluster-continuum model for accurate yet computationally efficient solvation modeling. Error cancellation between solvation energies of free and chelator-coordinated ions enables faster convergence, even with finite cluster sizes. Initial studies on the cryptates revealed consistent metal-ligand coordination patterns, with systematic variations influenced by ion size and charge, highlighting key structural features linked to binding selectivity. Further studies of a proprietary chelator have resulted in identification of previously unreported selectivity towards economically significant metals, which in-house experiments have confirmed, demonstrating the feasibility of this approach. By applying this methodology to new chelators targeting critical minerals such as lithium, cobalt, nickel and other strategic metals, we aim to accelerate the discovery of next-generation chelators for efficient recovery, recycling, and separation processes. This computational framework serves as the backbone of a high-throughput design pipeline tailored for sustainable resource utilization and may be applied to a wide range of systems to meet experimental needs.

computational materials↗

Meeting Report on the 3rd Chinese American Society for Mass Spectrometry Conference—Advancing Biological and Pharmaceutical Mass Spectrometry

Following the highly successful Chinese American Society for Mass Spectrometry (CASMS) conferences in the previous 2 years, the 3rd CASMS Conference was held virtually on August 28–31, 2023, using the Gather. Town platform to bring together scientists in the MS field. The conference offered a 4-day agenda with a scientific program consisting of two plenary lectures, and 14 parallel symposia in which a total of 70 speakers presented technological innovations and their applications in proteomics and biological MS and metabo-lipidomics and pharmaceutical MS. In addition, 16 invited speakers/panelists presented at two research-focused and three career development workshops. Moreover, 86 posters, 12 lightning talks, 3 sponsored workshops, and 11 exhibitions were presented, from which 9 poster awards and 2 lightning talk awards were selected. Furthermore, the conference featured four young investigator awardees to highlight early-career achievements in MS from our society. In conclusion, the conference provided a unique scientific platform for young scientists (i.e. graduate students, postdocs, and junior faculty/investigators) to present their research, meet with prominent scientists, learn about career development, and job opportunities (http://casms.org).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

PARIS: Predicting application resilience using machine learning

The traditional method to study application resilience to errors in HPC applications uses fault injection (FI), a time-consuming approach. Furthermore, while analytical models have been built to overcome the inefficiencies of FI, they lack accuracy. In this paper, we present PARIS, a machine-learning method to predict application resilience that avoids the time-consuming process of random FI and provides higher prediction accuracy than analytical models. PARIS captures the implicit relationship between application characteristics and application resilience, which is difficult to capture using most analytical models. We overcome many technical challenges for feature construction, extraction, and selection to use machine learning in our prediction approach. Our evaluation on 16 HPC benchmarks shows that PARIS achieves high prediction accuracy. PARIS is up to 450x faster than random FI (49x on average). Compared to the state-of-the-art analytical model, PARIS is at least 63% better in terms of accuracy and has comparable execution time on average.

97 MATHEMATICS AND COMPUTING↗

Efficient prediction of concentrating solar power plant productivity using data clustering

Concentrating solar power (CSP) plants convert solar energy to electricity and can be deployed with a thermal storage capability to shift electricity generation from time periods with available solar resource to those with high electricity demand or electricity price. Rigorous optimization of plant design and operational strategies can improve the market-competitiveness and commercial viability; however, such optimization may require hundreds of annual performance simulations, each of which can be computationally expensive when including considerations such as optimization of dispatch scheduling, sub-hourly time resolution, and stochastic effects due to uncertain weather or electricity price forecasts. This paper proposes a methodology to reduce the computational burden associated with simulation of electricity yield and revenue for CSP plants over a single- or multi-year period. Data-clustering techniques are employed to select a small number of limited-duration time blocks for simulation that, when appropriately weighted, can reproduce generation and revenue over a single year or within each year of a multi-year period. After selection of appropriate data features and weighting factors defining similarity between time-series profiles, the methodology captured annual revenue within 2.3%, 1.7%, or 1.2% using simulation of 10, 30, or 50 three-day exemplar time blocks, respectively, for each of three single-year location/weather/market scenarios and five plant configurations ranging from low to high solar multiple and storage capacity. When applied to multi-year datasets, the proposed methodology can capture inter-year variability that is unavailable from typical meteorological year (TMY) datasets while simultaneously requiring simulation of less than a single year of data.

14 SOLAR ENERGY↗

Prediction of Specificity of α-Conotoxins to Subtypes of Human Nicotinic Acetylcholine Receptors with Semi-supervised Machine Learning

Conotoxins are a family of highly toxic neurotoxins composed of cysteine-rich peptides produced by marine cone snails. The most lethal cone snail species to humans is Conus geographus, with fatality rates of up to ∼65% from a single sting, which is caused mostly by the activity of α-conotoxins against human nicotinic acetylcholine receptors (nAChRs). While sequence-based machine learning (ML) classifiers have been trained to identify targets of conotoxins binding voltage-gated ion channels, no ML model has been built to predict the subtype-specific nAChR targets of α-conotoxins. Here, we trained an ML model in a semi-supervised manner to predict the specificity of α-conotoxin binding toward different human nAChR subtypes to overcome the challenge of limited data in subtype-specific nAChR targets of α-conotoxins and the issue that one α-conotoxin can bind multiple nAChR subtypes with high selectivity. We considered additional features of sequences of α-conotoxins in training our ML model, including the secondary structure propensities and electrostatic properties, which resulted in better prediction capability for the ML model. Notably, we identify that most α-conotoxins bind to α3β2, α1γδ, and α7 subtypes of human nAChRs. Our findings from this study provide a framework for predicting targets of various kinds of toxins.

59 BASIC BIOLOGICAL SCIENCES↗

Enantioselective Allenoate-Claisen Rearrangement Using Chiral Phosphate Catalysts

Herein we report the first highly enantioselective allenoate-Claisen rearrangement using doubly axially chiral phosphate sodium salts as catalysts. This synthetic method provides access to β-amino acid derivatives with vicinal stereocenters in up to 95% ee. We also investigated the mechanism of enantioinduction by transition state (TS) computations with DFT as well as statistical modeling of the relationship between selectivity and the molecular features of both the catalyst and substrate. The mutual interactions of charge-separated regions in both the zwitterionic intermediate generated by reaction of an amine to the allenoate and the Na + -salt of the chiral phosphate leads to an orientation of the TS in the catalytic pocket that maximizes favorable noncovalent interactions. Crucial arene-arene interactions at the periphery of the catalyst lead to a differentiation of the TS diastereomers. Finally, these interactions were interrogated using DFT calculations and validated through statistical modeling of parameters describing noncovalent interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning-accelerated discovery of heat-resistant polysulfates for electrostatic energy storage

The development of heat-resistant dielectric polymers that withstand intense electric fields at high temperatures is critical for electrification. Balancing thermal stability and electrical insulation, however, is exceptionally challenging as these properties are often inversely correlated. A traditional intuition-driven polymer design approach results in a slow discovery loop that limits breakthroughs. Here we present a machine learning-driven strategy to rapidly identify high-performance, heat-resistant polymers. A trustworthy feed-forward neural network is trained to predict key proxy parameters and down select polymer candidates from a library of nearly 50,000 polysulfates. The highly efficient and modular sulfur fluoride exchange click chemistry enables successful synthesis and validation of selected candidates. A polysulfate featuring a 9,9-di(naphthalene)-fluorene repeat unit exhibits excellent thermal resilience and achieves ultrahigh discharged energy density with over 90% efficiency at 200 °C. Its exceptional cycling stability underscores its promise for applications in demanding electrified environments.

Li, He↗

A graph neural network-state predictive information bottleneck (GNN-SPIB) approach for learning molecular thermodynamics and kinetics

Molecular dynamics simulations offer detailed insights into atomic motions but face timescale limitations. Enhanced sampling methods have addressed these challenges but even with machine learning, they often rely on pre-selected expert-based features. Here, in this work, we present a Graph Neural Network-State Predictive Information Bottleneck (GNN-SPIB) framework, which combines graph neural networks and the state predictive information bottleneck to automatically learn low-dimensional representations directly from atomic coordinates. Tested on three benchmark systems, our approach predicts essential structural, thermodynamic and kinetic information for slow processes, demonstrating robustness across diverse systems. The method shows promise for complex systems, enabling effective enhanced sampling without requiring pre-defined reaction coordinates or input features.

Zou, Ziyue↗

Quantum model learning agent: characterisation of quantum systems through machine learning

Accurate models of real quantum systems are important for investigating their behaviour, yet are difficult to distil empirically. Here, we report an algorithm—the quantum model learning agent (QMLA)—to reverse engineer Hamiltonian descriptions of a target system. We test the performance of QMLA on a number of simulated experiments, demonstrating several mechanisms for the design of candidate Hamiltonian models and simultaneously entertaining numerous hypotheses about the nature of the physical interactions governing the system under study. QMLA is shown to identify the true model in the majority of instances, when provided with limited a priori information, and control of the experimental setup. Our protocol can explore Ising, Heisenberg and Hubbard families of models in parallel, reliably identifying the family which best describes the system dynamics. We demonstrate QMLA operating on large model spaces by incorporating a genetic algorithm to formulate new hypothetical models. The selection of models whose features propagate to the next generation is based upon an objective function inspired by the Elo rating scheme, typically used to rate competitors in games such as chess and football. In all instances, our protocol finds models that exhibit F 1 score ≥ 0.88 when compared with the true model, and it precisely identifies the true model in 72% of cases, whilst exploring a space of over 250 000 potential models. By testing which interactions actually occur in the target system, QMLA is a viable tool for both the exploration of fundamental physics and the characterisation and calibration of quantum devices.

97 MATHEMATICS AND COMPUTING↗

MIBiG 3.0: a community-driven effort to annotate experimentally validated biosynthetic gene clusters

Abstract With an ever-increasing amount of (meta)genomic data being deposited in sequence databases, (meta)genome mining for natural product biosynthetic pathways occupies a critical role in the discovery of novel pharmaceutical drugs, crop protection agents and biomaterials. The genes that encode these pathways are often organised into biosynthetic gene clusters (BGCs). In 2015, we defined the Minimum Information about a Biosynthetic Gene cluster (MIBiG): a standardised data format that describes the minimally required information to uniquely characterise a BGC. We simultaneously constructed an accompanying online database of BGCs, which has since been widely used by the community as a reference dataset for BGCs and was expanded to 2021 entries in 2019 (MIBiG 2.0). Here, we describe MIBiG 3.0, a database update comprising large-scale validation and re-annotation of existing entries and 661 new entries. Particular attention was paid to the annotation of compound structures and biological activities, as well as protein domain selectivities. Together, these new features keep the database up-to-date, and will provide new opportunities for the scientific community to use its freely available data, e.g. for the training of new machine learning models to predict sequence-structure-function relationships for diverse natural products. MIBiG 3.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

Improved precision in As speciation analysis with HERFD-XANES at the As K -edge: the case of As speciation in mine waste

High-energy-resolution fluorescence-detected (HERFD) X-ray absorption near-edge spectroscopy (XANES) is a spectroscopic method that allows for increased spectral feature resolution, and greater selectivity to decrease complex matrix effects compared with conventional XANES. XANES is an ideal tool for speciation of elements in solid-phase environmental samples. Accurate speciation of As in mine waste materials is important for understanding the mobility and toxicity of As in near-surface environments. In this study, linear combination fitting (LCF) was performed on synthetic spectra generated from mixtures of eight measured reference compounds for both HERFD-XANES and transmission-detected XANES to evaluate the improvement in quantitative speciation with HERFD-XANES spectra. The reference compounds arsenolite (As 2 O 3 ), orpiment (As 2 S 3 ), getchellite (AsSbS 3 ), arsenopyrite (FeAsS), kaňkite (FeAsO 4 ·3.5H 2 O), scorodite (FeAsO 4 ·2H 2 O), sodium arsenate (Na 3 AsO 4 ), and realgar (As 4 S 4 ) were selected for their importance in mine waste systems. Statistical methods of principal component analysis and target transformation were employed to determine whether HERFD improves identification of the components in a dataset of mixtures of reference compounds. LCF was performed on HERFD- and total fluorescence yield (TFY)-XANES spectra collected from mine waste samples. Arsenopyrite, arsenolite, orpiment, and sodium arsenate were more accurately identified in the synthetic HERFD-XANES spectra compared with the transmission-XANES spectra. In mine waste samples containing arsenopyrite and either scorodite or kaňkite, LCF with HERFD-XANES measurements resulted in fits with smaller R -factors than concurrently collected TFY measurements. The improved accuracy of HERFD-XANES analysis may provide enhanced delineation of As phases controlling biogeochemical reactions in mine wastes, contaminated soils, and remediation systems.

58 GEOSCIENCES↗

Electrical Measurement and Verification of Energy in DC Buildings

Today's selection of DC buildings features a diverse set of electrical topologies and turnkey solutions, and each has specific design trade-offs and optimizations. Designers desperately need standardized metrics and procedures for measurement and verification (M&V) to analyze and compare the advantages of each DC solution to traditional AC building networks. This work develops the Measurement-Informed Modeling (MIM) method, which can be used to determine full-building efficiency and energy savings. The MIM M&V procedure develops a building model, and refines the model with metered data. This work demonstrates the MIM method by measuring the full-building efficiency of two DC buildings operated by the Institute of Building Research in Shenzhen, China. The MIM procedure can ultimately be used to compare and improve the efficiency of various DC topologies.

buildings↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗