Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature vectors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Geoanalytical Evaluation of Saline Storage (GEESS) Geodatabase v2.0

The Geoanalytical Economic Evaluation of Saline Storage (GEESS) geodatabase was developed to support the United States Department of Energy (DOE) and National Energy Technology Laboratory (NETL) in their geologic carbon storage efforts by characterizing saline geologic formations present in the FECM/NETL CO2 Saline Storage Cost Model (CO2_S_COM) [1]. Using publicly available literature and data, the GEESS geodatabase characterizes 57 geologic formations across the lower-48 U.S. states in what are called Fully Integrated Geodatabases (FIGs). The FIG is a vector polygon feature containing thousands or tens of thousands of individual polygons, which each contain discrete geologic parameter values. A list of the critical geologic parameters that are characterized in the GEESS geodatabase are described in the “Processing Steps and Workflow” part of the ReadMe file, as well as the Data Catalog accompanying the GEESS geodatabase. The FIG is the basis of the GEESS geodatabase and is the direct representation of the collected geologic data. In addition to the FIG, the GEESS system contains grid files. Due to the complexity of the FIGs, grids are used to sample the geologic data so they can be exercised within CO2_S_COM. The grid files contain the geologic data sampled from the FIG, as well as estimates of “Plume Uncertainty Diameter” and “First-year Break-even Price of CO2” derived from CO2_S_COM based on the GEESS grid data.

carbon↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

Non-volatile electric control of antiferromagnetic states on nanosecond timescales

Electrical manipulation of antiferromagnetic (AFM) states, a cornerstone of AFM spintronics, is a great challenge, requiring novel material platforms. Here we report the full control over AFM states by voltage pulses in the insulating Co 3 O 4 spinel well below its Néel temperature. We show that the strong linear magnetoelectric effect is fully governed by the orientation of the Néel vector. As a unique feature of Co 3 O 4 , the magnetoelectric energy can easily overcome the weak magnetocrystalline anisotropy, thus, the N´eel vector can be manipulated on demand, either rotated smoothly or reversed suddenly, by combined electric and magnetic fields. We achieve the non-volatile switching within a few tens of nanoseconds between time-reversed AFM states in macroscopic volumes by voltage pulses. These observations render quasi-cubic antiferromagnets, like Co 3 O 4 , an ideal platform for the ultrafast (pico- to nanosecond) manipulation of microscopic AFM domains and may pave the way for the realization of AFM spintronic devices.

36 MATERIALS SCIENCE↗

Spatially resolved polarization swings in the supermassive binary black hole candidate OJ 287 with first Event Horizon Telescope observations

We present the first Event Horizon Telescope 1.3 mm observations of the supermassive binary black hole candidate OJ 287. The observations achieved an unprecedented angular resolution of 18 μas and reveal significant structural and polarization variability over just five days, marking the shortest timescale on which such changes have been directly imaged in this source. The inner jet exhibits a twisted ridgeline structure, with features displaying apparent superluminal motions up to about 22 c. The linear polarization maps reveal three main polarized features whose electric-vector position angles (EVPAs) change substantially over the time span of our observations, including a component with a radial polarization consistent with being produced by a recollimation shock. Most notably, we directly resolved two innermost jet components whose EVPAs rotate in opposite directions. The faster component, moving at 2.4 ± 0.9 μas/day (17.4 ± 6.5 c), exhibits counterclockwise EVPA swings of roughly 3.7° per day, while the slower component, with a proper motion of 1.4 ± 0.3 μas/day (10.2 ± 2.2 c), rotates clockwise at approximately 2.5° per day. Previous studies inferred helical magnetic fields in AGN jets from time-resolved or integrated polarization variability but lacked the angular resolution to directly image this effect. Our results provide spatially resolved evidence that a helical magnetic field threads the jet’s collimation and acceleration zone, ruling out models based on the superposition of unresolved components. Our analysis suggests that propagating shocks interact with a Kelvin–Helmholtz plasma instability, illuminating different phases of the helical magnetic field and producing the observed polarization spatial and temporal variability. Moreover, our model naturally accounts for the more rapid polarization rotation observed in the faster moving component. Our model predicts even more rapid swings in polarization, which could be tested with future observations featuring a more densely sampled time coverage.

OJ 287↗

Dynamic Graph Sequence Data from Simulated Neutron Reflectometry Measurements

This dataset comprises dynamic graph sequences derived from simulated in-situ neutron reflectometry measurements, capturing the gradual evolution of a layer structure over time. Each graph sequence represents a synthetic sample, with node features detailing the scattering vector and corresponding reflectivity measurements, while adjacency matrices have corresponding reference material parameters attached as metadata. The dataset spans multiple sets, each with a different number of sequences, offering a comprehensive basis for training models that handle dynamic input sequences with embedded physics. This dataset is particularly suited for tackling inverse problems with hidden physical states that evolve over time, challenges that are typically difficult to address using conventional iterative fitting methods.

36 MATERIALS SCIENCE↗

RT-EZ: A Golden Gate Assembly Toolkit for Streamlined Genetic Engineering of Rhodotorula toruloides

For economic and sustainable biomanufacturing, the oleaginous yeast Rhodotorula toruloides has emerged as a promising platform for producing biofuels, pharmaceuticals, and other valuable chemicals. However, genetic manipulation of R. toruloides has been limited by its high GC content and the lack of a replicating plasmid, necessitating gene integration into the genome of the yeast. To address these challenges, we developed the RT-EZ (R. toruloides Efficient Zipper) toolkit, a versatile tool based on Golden Gate assembly, designed to streamline R. toruloides engineering with improved efficiency and flexibility. The RT-EZ toolkit simplifies vector construction by incorporating new features such as bidirectional promoters and 2A peptides, color-based screening using RFP, and sequences optimized for both Agrobacterium tumefaciens-mediated transformation (ATMT) and easy linearization, enabling straightforward selection and transformation. Notably, the RT-EZ kit can be used to construct an expression cassette with four different genes in one assembly reaction, significantly improving vector construction speed and efficiency. The utility of the RT-EZ toolkit was demonstrated through the successful synthesis of arachidonic acid in R. toruloides by coexpressing fatty acid elongases and desaturases. Furthermore, this result underscores the potential of the RT-EZ toolkit to advance synthetic biology in R. toruloides, providing a streamlined method for addressing genetic engineering challenges in the yeast.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data for "RT-EZ: A Golden Gate Assembly Toolkit for Streamlined Genetic Engineering of Rhodotorula toruloides"

For economic and sustainable biomanufacturing, the oleaginous yeast Rhodotorula toruloides has emerged as a promising platform for producing biofuels, pharmaceuticals, and other valuable chemicals. However, genetic manipulation of R. toruloides has been limited by its high GC content and the lack of a replicating plasmid, necessitating gene integration into the genome of the yeast. To address these challenges, we developed the RT-EZ ( R. toruloides Efficient Zipper) toolkit, a versatile tool based on Golden Gate assembly, designed to streamline R. toruloides engineering with improved efficiency and flexibility. The RT-EZ toolkit simplifies vector construction by incorporating new features such as bidirectional promoters and 2A peptides, color-based screening using RFP, and sequences optimized for both Agrobacterium tumefaciens-mediated transformation (ATMT) and easy linearization, enabling straightforward selection and transformation. Notably, the RT-EZ kit can be used to construct an expression cassette with four different genes in one assembly reaction, significantly improving vector construction speed and efficiency. The utility of the RT-EZ toolkit was demonstrated through the successful synthesis of arachidonic acid in R. toruloides by coexpressing fatty acid elongases and desaturases. This result underscores the potential of the RT-EZ toolkit to advance synthetic biology in R. toruloides , providing a streamlined method for addressing genetic engineering challenges in the yeast.

gene editing↗

A Geodatabase Designed to Inform and Support Safe CO2 Transport-Route Planning

The National Energy Technology Laboratory developed the Carbon Capture and Storage (CCS) Pipeline Route Planning Database to inform safe and sustainable CO2 transport planning in support of decarbonization efforts. This comprehensive Esri Geodatabase contains over 90 GBs of data and 60+ spatial layers representing key considerations including natural hazards, infrastructure, energy, and social justice. Leveraging ArcGIS Pro, nationwide raster, and vector datasets containing millions of features were processed to be easily digestible for users and complex modeling software.

Romeo, Lucy↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation↗

Low-energy electronic structure in the unconventional charge-ordered state of ScV6Sn6

Abstract Kagome vanadates A V 3 Sb 5 display unusual low-temperature electronic properties including charge density waves (CDW), whose microscopic origin remains unsettled. Recently, CDW order has been discovered in a new material ScV 6 Sn 6 , providing an opportunity to explore whether the onset of CDW leads to unusual electronic properties. Here, we study this question using angle-resolved photoemission spectroscopy (ARPES) and scanning tunneling microscopy (STM). The ARPES measurements show minimal changes to the electronic structure after the onset of CDW. However, STM quasiparticle interference (QPI) measurements show strong dispersing features related to the CDW ordering vectors. A plausible explanation is the presence of a strong momentum-dependent scattering potential peaked at the CDW wavevector, associated with the existence of competing CDW instabilities. Our STM results further indicate that the bands most affected by the CDW are near vHS, analogous to the case of A V 3 Sb 5 despite very different CDW wavevectors.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Observation of band splitting and magnetically induced band structure reconstruction in TbTi 3 ⁢Bi 4

The magnetic kagome materials are a promising platform to study the interplay between magnetism, topology, and correlated electronic phenomena. Among these materials, the 𝑅⁢Ti 3 ⁢Bi 4 family received a great deal of attention recently because of its chemical versatility and wide range of magnetic properties. Here, we use angle-resolved photoemission spectroscopy measurements and density functional theory calculations to investigate the electronic structure of TbTi 3 ⁢Bi 4 in paramagnetic and antiferromagnetic phases. Our experimental results show the presence of unidirectional band splitting of unknown nature in both phases. In addition, we observed a complex reconstruction of the band structure in the antiferromagnetic phase. Furthermore, some aspects of this reconstruction are consistent with effects of additional periodicity introduced by the magnetic ordering vector, while the nature of several other features remains unknown.

Angle-resolved photoemission spectroscopy↗

Leveraging intermediate resonances to probe CP violation at colliders

We explore the phenomenological impact of interference in tree-level contributions to three-body final states in $2\rightarrow 3$ scattering processes. This work introduces a novel search strategy leveraging asymmetries to enable sensitivity to CP-violating effects in less well-explored regions of phase space. Analytically, we demonstrate the effectiveness of this observable in probing interference between Standard Model charged-current decays and effective left-handed vector interactions, illustrated in a toy model featuring a scalar leptoquark, $S_1 \sim (3, 1, -\,1/3)$. Numerically, we apply this framework to studying the process $pp\rightarrow b \tau \nu $; unlike traditional high-$p_T$ searches or “bump hunts”, this approach utilizes an intermediate energy regime – where new physics is neither light enough to be produced on shell or heavy enough to justify an effective field theory treatment. A proof-of-principle analysis at parton level demonstrates a percent-level asymmetry, with sensitivity also to BSM weak-CP phase. While the specific phase sensitivity is diminished at particle level due to showering and detector effects, a machine learning classifier can recover sensitively to the presence of SM-BSM interference, significantly outperforming standard analysis methods. Notably discrimination between BSM signal and SM background could be achieved at the 2$\sigma $ level for the current LHC dataset and 8$\sigma $ at the High-Luminosity LHC. Moreover, this asymmetry observable as defined can also be more broadly applied to other searches for CP-violation in $2\rightarrow 3$ processes in present and future collider environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

A Data-Driven Approach for High-Impedance Fault Localization in Distribution Systems

Accurate and quick identification of high-impedance faults (HIFs) is critical for the reliable operation of distribution systems. Unlike other faults in power grids, HIFs are very difficult to detect by conventional overcurrent relays due to the low fault current. Although HIFs can be affected by various factors, the voltage-current characteristics can substantially imply how the system responds to the disturbance and thus provides opportunities to effectively localize HIFs. In this work, we propose a data-driven approach for the identification of HIF events. To tackle the nonlinearity of the voltage-current trajectory, first, we formulate optimization problems to approximate the trajectory with piecewise functions. Then we collect the function features of all segments as inputs and use the support vector machine approach to efficiently identify HIFs at different locations. Numerical studies on the IEEE 123-node test feeder demonstrate the validity and accuracy of the proposed approach for real-time HIF identification.

explainable artificial intelligence↗