Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “embedded methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Online learning of quadratic manifolds from streaming data for nonlinear dimensionality reduction and nonlinear model reduction

Here, this work introduces an online greedy method for constructing quadratic manifolds from streaming data, designed to enable in situ analysis of numerical simulation data on the Petabyte scale. Unlike traditional batch methods, which require all data to be available upfront and take multiple passes over the data, the proposed online greedy method incrementally updates quadratic manifolds in one pass as data points are received, eliminating the need for expensive disk input/output operations as well as storing and loading data points once they have been processed. A range of numerical examples demonstrate that the online greedy method learns accurate quadratic manifold embeddings while being capable of processing data that far exceed common disk input/output capabilities and volumes as well as main-memory sizes.

97 MATHEMATICS AND COMPUTING↗

Heliostat Sizing Methodology for Solar Heat for Industrial Processes

This study presents a method to obtain a heliostat size that minimizes the levelized cost of a heliostat-based concentrating solar thermal system for industrial process heat (IPH) applications at operating temperatures from 565 to 1550 degrees Celsius. The method extends prior work by embedding a routine for system design that obtains near-optimal subsystem sizes, increasing the fidelity of drive cost functions, and adding an optical performance model. An illustrative business case is developed for Daggett, California, targeting specified annual thermal energy outputs of 50 to 400 GWhth. Optical performance is modeled using verified estimates from the literature. A surrogate heliostat cost model, derived from commercial heliostat designs and scaled for production volume, installation, and operations and maintenance costs, is used to develop cost functions. Results show that heliostat size strongly affects the levelized cost of heat (LCOH), producing a characteristic U-shaped trend with a robust near-optimal window of 8 - 12 m2; the heliostat size producing the lowest project cost in our study grows slightly as the project size increases, and is reduced as the operating temperature increases. The findings in this study are consistent with the general trend of smaller heliostats under deployment at existing projects for high-temperature industrial process heat and reflect the significant reduction in power electronics and other per-heliostat costs. The methodology we propose is general and can be tailored to revised cost curves as the technology continues to evolve.

14 SOLAR ENERGY↗

DOME: Directional medical embedding vectors from Electronic Health Records

Motivation: The increasing availability of Electronic Health Record (EHR) systems has created enormous potential for translational research. Recent developments in representation learning techniques have led to effective large-scale representations of EHR concepts along with knowledge graphs that empower downstream EHR studies. However, most existing methods require training with patient-level data, limiting their abilities to expand the training with multi-institutional EHR data. On the other hand, scalable approaches that only require summary-level data do not incorporate temporal dependencies between concepts. Methods: We introduce a DirectiOnal Medical Embedding (DOME) algorithm to encode temporally directional relationships between medical concepts, using summary-level EHR data. Specifically, DOME first aggregates patient-level EHR data into an asymmetric co-occurrence matrix. Then it computes two Positive Pointwise Mutual Information (PPMI) matrices to correspondingly encode the pairwise prior and posterior dependencies between medical concepts. Following that, a joint matrix factorization is performed on the two PPMI matrices, which results in three vectors for each concept: a semantic embedding and two directional context embeddings. They collectively provide a comprehensive depiction of the temporal relationship between EHR concepts. Results: We highlight the advantages and translational potential of DOME through three sets of validation studies. First, DOME consistently improves existing direction-agnostic embedding vectors for disease risk prediction in several diseases, for example achieving a relative gain of 5.5% in the area under the receiver operating characteristic (AUROC) for lung cancer. Second, DOME excels in directional drug-disease relationship inference by successfully differentiating between drug side effects and indications, correspondingly achieving relative AUROC gain over the state-of-the-art methods by 10.8% and 6.6%. Finally, DOME effectively constructs directional knowledge graphs, which distinguish disease risk factors from comorbidities, thereby revealing disease progression trajectories. The source codes are provided at https://github.com/celehs/Directional-EHRembedding.

60 APPLIED LIFE SCIENCES↗

Towards utility-scale electronic structure with sample-based quantum bootstrap embedding

One of the main applications for which quantum computers are hoped to find utility is in simulating ground state energies and other observables of molecular chemical systems. The recently proposed sample-based diagonalization method is a readily implementable method for this task on current-day hardware using short circuit depths and has been demonstrated on as many as 85 qubits in recent studies. In this work, we combine the recently proposed quantum bootstrap embedding (QBE) method with sampled-based diagonalization (QBE-SQD) and present the first benchmarking study of the QBE method on real quantum hardware, ibm_pittsburgh, a Heron r3 processor with 156 qubits. Our test system is a hydrogen ring with 8 hydrogen atoms in the cc-pVDZ basis. We show that for this system, QBE-SQD using an active space of (8e, 19o) per fragment with a 43 qubit footprint produces a ground state energy accuracy which exceeds that of an SQD calculation with an (8e, 30o) active space with a 67 qubit footprint when using a comparable number of Slater determinants. This demonstrates that the use of quantum bootstrap embedding techniques is a promising path towards extending the capabilities of state-of-the-art quantum eigensolvers on near-term devices.

Bierman, Joel [North Carolina State University, Ra↗

Chemical templates of the Central Molecular Zone

Context . The Central Molecular Zone (CMZ) of the Milky Way exhibits extreme conditions, including high gas densities, elevated temperatures, enhanced cosmic-ray ionization rates, and large-scale dynamics. This makes it a perfect laboratory for astrochemical studies. With large-scale molecular surveys revealing increasing chemical and physical complexity in the CMZ, it is essential to develop robust methods to decode the chemical information embedded in this extreme region. Aims . A key step to interpreting the molecular richness found in the CMZ is building chemical templates tailored to its diverse conditions. In particular, understanding how CMZ environments affect shock and protostellar chemistry is crucial. The combined impact of high ionization, elevated temperatures, and dense gas remains insufficiently explored for observable tracers. Methods . For this study, we utilized UCLCHEM , a gas-grain time-dependent chemical model, to link physical conditions with their corresponding molecular signatures and identify key tracers of temperature, density, ionization, and shock activity. To achieve this, we ran a grid of models of shocks and protostellar objects representative of typical CMZ conditions, focusing on 24 species, including complex organic molecules. Results . Shocked and protostellar environments show distinct evolutionary timescales (≲10 4 vs. ≳10 4 years); 300 K emerges as a key temperature threshold for chemical differentiation. We find that cosmic-ray ionization and temperature are the main drivers of chemical trends. HCO + , H 2 CO, and CH 3 SH trace ionization, while HCO, HCO + , CH 3 SH, CH 3 NCO, and HCOOCH 3 show consistent abundance contrasts between shocks and protostellar regions over similar temperature ranges. Conclusions . We characterized the behavior of 24 species in protostellar and shock-related environments. While our models underpredict some complex organics in shocks, they reproduce observed trends for most species, supporting scenarios involving a need for recurring shocks in Galactic Center clouds and enhanced ionization toward Sgr B2(N2). Future work should assess the role of shock recurrence and metallicity in shaping chemistry.

Galaxy: center↗

Divertor detachment and heat exhaust mitigation control in KSTAR with tungsten divertor

KSTAR has recently undergone an upgrade to use a new tungsten divertor to run experiments in ITER-relevant scenarios. Even with a high melting point of tungsten, it is important to control the heat flux impinging on tungsten divertor targets to minimize sputtering and contamination of the core plasma. Heat flux on the divertor is often controlled by increasing the degree of detachment of scrape-off layer plasma from the target plates. In this work, we have demonstrated successful divertor detachment and heat exhaust dissipation control experiments using two different methods. The first method uses attachment fraction as a control variable which is estimated using ion saturation current measurements from embedded Langmuir probes in the divertor. The second method uses a novel machine-learning-based surrogate model of 2D UEDGE simulation database, DivControlNN. We demonstrated running inference operation of DivControlNN in realtime to estimate heat flux at the divertor and use it as the control variable in a feedback loop with impurity gas flow. We present interesting insights from these experiments including a systematic approach to tuning controllers and discuss future improvements in the control infrastructure and control variables for future burning plasma experiments.

KSTAR tungsten divertor operations↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Latent Catalysis as a Platform for Accessing Diverse Material Properties in Vat Photopolymerization 3D Printing

Vat photopolymerization (VP) 3D printing is an attractive strategy to manufacture customized polymer parts. The properties of printed materials are limited by the need to employ a low viscosity liquid resin and achieve rapid polymerization kinetics. To circumvent this limitation, dual‐cure methods have been developed using reagents embedded in the liquid resin formulation; however, the reagent‐based approach requires the discovery and optimization of new chemistry for each desired material. Here, in this work, we demonstrate a catalytic, dual‐cure platform that enables access to both Nylon‐6 and polyester interpenetrating networks through VP 3D printing under a universal approach. Structure–reactivity relationships of the latent NHC catalysts led to the identification of a magnesium chloride–NHC adduct as a latent catalyst that is orthogonal to radical polymerization and can be unmasked at elevated temperatures post‐printing to initiate ring‐opening polymerization of lactones and lactams. This strategy results in access to semicrystalline materials, which are a challenging morphology to access via VP 3D printing, that have attractive mechanical properties and can be printed at high resolution. This work represents the first photochemical‐based 3D printing of Nylon‐based materials and demonstrates the value of catalytic approaches to access new material properties in VP 3D printing.

Colliver, Cali N. [University of North Carolina, C↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Methods for Color Center Preserving Hydrogen‐Termination of Diamond

Abstract Chemical functionalization of diamond surfaces by hydrogen is an important method for controlling the charge state of near‐surface fluorescent color centers, an essential process in fabricating devices such as diamond field‐effect transistors and chemical sensors, and a required first step for realizing families of more complex terminations through subsequent chemical processing. In all these cases, termination is typically achieved using hydrogen plasma sources that can etch or damage the diamond, as well as deposited materials or embedded color centers. This work explores alternative methods for lower‐damage hydrogenation of diamond surfaces, specifically the annealing of diamond samples in high‐purity, non‐explosive mixtures of nitrogen and hydrogen gas, and the exposure of samples to microwave hydrogen plasmas in the absence of intentional stage heating. The effectiveness of these methods are characterized by x‐ray photoelectron spectroscopy (XPS), and comparison of the results to density‐functional modelling of the surface hydrogenation energetics implicates surface oxygen ligands as the primary factor limiting the termination quality of annealed samples. Finally, photoluminescence (PL) spectroscopy is used to verify that both the annealing and reduced sample temperature plasma methods are non‐destructive to near‐surface ensembles of nitrogen‐vacancy (NV) centers, in stark contrast to plasma treatments that use heated sample stages.

36 MATERIALS SCIENCE↗

Accelerating Embedding Potential Optimization by Reconstructing the Pseudo-Valence Electron Density

Density functional embedding theory (DFET) enables use of electronic structure methods with higher accuracy than density functional theory in a local region, with applications thus far ranging from (photo/electro)catalysis to reactions in solution. DFET partitions a large collection of atoms into smaller groups that interact via a shared embedding (interaction) potential V emb , determined via functional optimization. The optimized effective potential (OEP) process used to optimize V emb is time-consuming and becomes a computational bottleneck due to sharp, oscillating features of V emb near nuclei. Here, similar to pseudopotential theory, by reconstructing electron densities used in the OEP process from smoother pseudo-valence-only (PVO) electron densities as proxies for total densities of the full system and subsystems, we can retain accuracy in the embedded electronic structure calculations while potentially reducing the overhead of V emb construction, within the projector augmented-wave (PAW) formalism. We explore three different chemical reactions as exemplars to test PVO–DFET, namely, H 2 dissociative adsorption on a Cu(111) surface, H 2 O adsorption on a Pt(111) surface, and aqueous [Ca 2+ –SO 4 2– ] ion-pair formation. The PVO approximation works well for all three systems with minimal loss of accuracy (∼10–70 meV error relative to the original exact-derivative (ED) approach) while accelerating V emb generation for the Cu and Pt systems respectively by 20× and 5×. Given proper numerical convergence parameters, the spatial distributions of differences between PVO- and ED-based V emb outside the core regions are small, explaining the exceptional agreement between the two approaches. Finally, we anticipate that this more efficient PVO–DFET approximation will be useful whenever computation of V emb is much more expensive than subsequent embedded high-level electron correlation calculations.

approximation↗

Visualizing Temporal Topic Embeddings with a Compass

—Dynamic topic modeling is useful at discovering the development and change in latent topics over time. However, present methodology relies on algorithms that separate document and word representations. This prevents the creation of a meaningful embedding space where changes in word usage and documents can be directly analyzed in a temporal context. This paper proposes an expansion of the compass-aligned temporal Word2Vec methodology into dynamic topic modeling. Such a method allows for the direct comparison of word and document embeddings across time in dynamic topics. This enables the creation of visualizations that incorporate temporal word embeddings within the context of documents into topic visualizations. In experiments against the current state-of-the-art, our proposed method demonstrates overall competitive performance in topic relevancy and diversity across temporal datasets of varying size. Simultaneously, it provides insightful visualizations focused on temporal word embeddings while maintaining the insights provided by global topic evolution, advancing our understanding of how topics evolve over time.

Cluster analysis↗

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

Embedded Aluminum Nitride Sensors for Advanced Reactors

This project aims to manipulate the of growth of aluminum nitride (AlN) inclusions in an iron-chromium-aluminum (FeCrAl) alloy, using the principle of powder metallurgy and heat treatment methods. to promote the formation of AlN phase within FeCrAl for embedded sensing. This effort seeks to fill in a gap with respect to robust sensor hardware for ubiquitous structural health monitoring of advanced nuclear reactors. A systematic evaluation of various solid-state methods will be performed to understand the influence of process conditions on AlN growth and to promote the growth of desirable AlN phase. Fabricated specimens will be analyzed to investigate the AlN structures that are formed using a suite of tools to characterize the concentration, morphology, and distribution of AlN within the FeCrAl substrate.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Polymer-fiber-reinforced polymers with enhanced interfacial bonding between polypropylene fiber and polyethylene matrix

Self-reinforced composites (SRCs) consist of reinforcing fibers and a base matrix made of the same thermoplastic polymer, offering lightweight, recyclability, and sustainability benefits. However, limited research exists on composites where the reinforcing thermoplastic polymer fibers differ from the base thermoplastic matrix. Here, this study focuses on investigating the mechanical behavior of such composites and exploring different surface modification methods to enhance the fiber/matrix interfacial bonding using polypropylene fibers and a polyethylene matrix as an example. It is shown that surface treatment with a commercial adhesion promoter containing n-butyl acetate significantly improves the interfacial shear strength between polypropylene fibers and the polyethylene matrix, increasing it by 145% compared to other methods investigated. Additionally, increasing the length of the embedded polymer fiber in the matrix leads to a notable increase in specific interfacial energy. Consequently, the thermoplastic polymer-fiber-reinforced polymers (PFRPs) using surface-treated woven polypropylene fabrics and a polyethylene matrix exhibit a 20% higher tensile strength and a 65% higher toughness compared to non-treated PFRPs. This study also shows that specific mechanical properties (normalized by the composite density) of the investigated woven PFRPs are similar to those of non-treated SRCs under uni-axial tension. Particularly, their ductility outperforms carbon-/glass-/aramid-fiber-reinforced polymers by at least 6 times at a same fiber volume fraction. The investigation of such composites and the exploration of surface modification methods present important progress in the field of thermoplastic PFRPs, which serve as a solution for addressing concerns related to recyclability and sustainability.

Fiber pull-out↗

Replacing non-biomedical concepts improves embedding of biomedical concepts

Embeddings are semantically meaningful representations of words in a vector space, commonly used to enhance downstream machine learning applications. Traditional biomedical embedding techniques often replace all synonymous words representing biological or medical concepts with a unique token, ensuring consistent representation and improving embedding quality. However, the potential impact of replacing non-biomedical concept synonyms has received less attention. Embedding approaches often employ concept replacement to replace concepts that span multiple words, such as non-small-cell lung carcinoma, with a single concept identifier (e.g., D002289). Also, all synonyms of each concept are merged into the same identifier. Here, we additionally leveraged WordNet to identify and replace sets of non-biomedical synonyms with their most common representatives. This combined approach aimed to reduce embedding noise from non-biomedical terms while preserving the integrity of biomedical concept representations. We applied this method to 1,055 biomedical concept sets representing molecular signatures or medical categories and assessed the mean pairwise distance of embeddings with and without non-biomedical synonym replacement. A smaller mean pairwise distance was interpreted as greater intra-cluster coherence and higher embedding quality. Embeddings were generated using the Word2Vec algorithm applied to a corpus of 10 million PubMed abstracts. Our results demonstrate that the addition of non-biomedical synonym replacement reduced the mean intra-cluster distance by an average of 8%, suggesting that this complementary approach enhances embedding quality. Future work will assess its applicability to other embedding techniques and downstream tasks. Python code implementing this method is provided under an open-source license.

algorithms↗

Projector‐Based Quantum Embedding Study of Iron Complexes

Projection‐based embedding theory (PBET) is used to calculate and assess the challenging spin‐crossover energies for a selection of small Fe‐containing systems by embedding the metal center into the frozen potential of the ligands. MP2, CCSD, and CCSD(T) are embedded in potentials from the SCAN and r 2 SCAN functionals and compared with the canonical values for the constituent methods and previously reported reference values. Considering the PBET calculations as a correction for the underlying DFT, the embedding calculations are able to provided improvement for most cases. In some cases, the PBET methods are able to compensate for limitations in the wave function methods and produce results similar to more rigorous calculations from the literature. For the systems with spin‐crossover energies near zero, the current methodology fails to provide consistent improvement. In conclusion, the isolated recalculation of the electronic structure around the metal center when embedded into a DFT treatment of the ligand field shows promise as a pragmatic and lower cost treatment compared to the canonical treatment of the whole system of the difficult class of spin‐crossover complexes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗