Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Compositions and methods for engineering oil content in plants

Compositions and methods for producing plants with enhanced oil content and higher seed yield are disclosed. The transgenic plant comprises a polynucleotide encoding a monoacylglycerol O-acyltransferase 1 (MGAT1) operatively linked to a plant-expressible promoter; a polynucleotide encoding a phosphatidylcholine diacylglycerol cholinephosphotransferase 1 (PDCT1) operatively linked to a plant-expressible promoter; a polynucleotide encoding a suppressor of expression of Sugar Dependent 1 (SPD1) operatively linked to a plant-expressible promoter; a polynucleotide encoding a diacylglyerol acyltransferase (DGAT1) operatively linked to a plant-expressible promoter and a polynucleotide encoding a glycerol-3-phosphate dehydrogenase (GPD1) operatively linked to a plant-expressible promoter; or a combination thereof.

Dhankher, Om Parkash↗

Efficient Floating-Point Arithmetic on Fault-Tolerant Quantum Computers

We propose a novel floating-point encoding scheme that builds on prior work involving fixed-point encodings. We encode floating-point numbers using Two's Complement fixed-point mantissas and Two's Complement integral exponents. We used our proposed approach to develop quantum algorithms for fundamental arithmetic operations, such as bit-shifting, reciprocation, multiplication, and addition. We prototyped and investigated the performance of the floating-point encoding scheme on quantum computer simulations by performing reciprocation on randomly drawn inputs and by solving first-order ordinary differential equations, while varying the number of qubits in the encoding. We observed rapid convergence to the exact solutions as we increased the number of qubits and a significant reduction in the number of ancilla qubits required for reciprocation when compared with similar approaches.

Serrallés, José Cruz [Weill Cornell Med. Coll.]↗

When does global attention help: a unified empirical study on atomistic graph learning

Graph neural networks (GNNs) are widely used as surrogates for costly experiments and first-principles simulations to study the behavior of compounds at atomistic scale, and their architectural complexity is constantly increasing to enable the modeling of complex physics. While most recent GNNs combine more traditional message passing neural networks (MPNNs) layers to model short-range interactions with more advanced graph transformers (GTs) with global attention mechanisms to model long-range interactions, it is still unclear when global attention mechanisms provide real benefits over well-tuned MPNN layers due to inconsistent implementations, features, or hyperparameter tuning. We introduce the first unified, reproducible benchmarking framework–built on HydraGNN–that enables seamless switching among four controlled model classes: MPNN, MPNN with chemistry/topology encoders, GPS-style hybrids of MPNN with global attention, and fully fused localglobal models with encoders. Using seven diverse open-source datasets for benchmarking across regression and classification tasks, we systematically isolate the contributions of message passing, global attention, and encoder-based feature augmentation. Our study shows that encoder-augmented MPNNs form a robust baseline, while fused localglobal models yield the clearest benefits for properties governed by long-range interaction effects. We further quantify the accuracycompute trade-offs of attention, reporting its overhead in memory. Together, these results establish the first controlled evaluation of global attention in atomistic graph learning and provide a reproducible testbed for future model development.

Equivariant graph neural networks↗

Metagenomic strategies identify diverse integron–integrase and antibiotic resistance genes in the Antarctic environment

The objective of this study is to identify and analyze integrons and antibiotic resistance genes (ARGs) in samples collected from diverse sites in terrestrial Antarctica. Integrons were studied using two independent methods. One involved the construction and analysis of intI gene amplicon libraries. In addition, we sequenced 17 metagenomes of microbial mats and soil by high-throughput sequencing and analyzed these data using the IntegronFinder program. As expected, the metagenomic analysis allowed for the identification of novel predicted intI integrases and gene cassettes (GCs), which mostly encode unknown functions. However, some intI genes are similar to sequences previously identified by amplicon library analysis in soil samples collected from non-Antarctic sites. ARGs were analyzed in the metagenomes using ABRIcate with CARD database and verified if these genes could be classified as GCs by IntegronFinder. We identified 53 ARGs in 15 metagenomes, but only four were classified as GCs, one in MTG12 metagenome (Continental Antarctica), encoding an aminoglycoside-modifying enzyme (AAC(6´)acetyltransferase) and the other three in CS1 metagenome (Maritime Antarctica). One of these genes encodes a class D β-lactamase (blaOXA-205) and the other two are located in the same contig. One is part of a gene encoding the first 76 amino acids of aminoglycoside adenyltransferase (aadA6), and the other is a qacG2 gene.

59 BASIC BIOLOGICAL SCIENCES↗

Novel pore size-controlled, susceptibility matched, 3D-printed MRI phantoms

We report the design concept and fabrication of MRI phantoms, containing blocks of aligned microcapillaires that can be stacked into larger arrays to construct diameter distribution phantoms or fractured, to create a “powder-averaged” emulsion of randomly oriented blocks for vetting or calibrating advanced MRI methods, that is, diffusion tensor imaging, AxCaliber MRI, MAP-MRI, and multiple pulsed field gradient or double diffusion-encoded microstructure imaging methods. The goal was to create a susceptibility-matched microscopically anisotropic but macroscopically isotropic phantom with a ground truth diameter that could be used to vet advanced diffusion methods for diameter determination in fibrous tissues. Two-photon polymerization, a novel three-dimensional printing method is used to fabricate blocks of capillaries. Double diffusion encoding methods were employed and analyzed to estimate the expected MRI diameter. Susceptibility-matched microcapillary blocks or modules that can be assembled into large-scale MRI phantoms have been fabricated and measured using advanced diffusion methods, resulting in microscopic anisotropy and random orientation. This phantom can vet and calibrate various advanced MRI methods and multiple pulsed field gradient or diffusion-encoded microstructure imaging methods. We demonstrated that two double diffusion encoding methods underestimated the ground truth diameter.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

Robust deep learning framework for constitutive relations modeling

Modeling the full-range deformation behaviors of materials under complex loading and materials conditions is a significant challenge for constitutive relations (CRs) modeling. Here, we propose a general encoder-decoder deep learning framework that can model high-dimensional stress-strain data and complex loading histories with robustness and universal capability. The framework employs an encoder to project high-dimensional input information (e.g., loading history, loading conditions, and materials information) to a lower-dimensional hidden space and a decoder to map the hidden representation to the stress of interest. We evaluated various encoder architectures, including gated recurrent unit (GRU), GRU with attention, temporal convolutional network (TCN), and the Transformer encoder, on two complex stress-strain datasets that were designed to include a wide range of complex loading histories and loading conditions. All architectures achieved excellent test results with an root-mean-square error (RMSE) below 1 MPa. Additionally, we analyzed the capability of the different architectures to make predictions on out-of-domain applications, with an uncertainty estimation based on deep ensembles. The proposed approach provides a robust alternative to empirical/semi-empirical models for CRs modeling, offering the potential for more accurate and efficient materials design and optimization.

36 MATERIALS SCIENCE↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Latent code-based fusion: A Volterra neural network approach

We propose a deep structure encoder using Volterra Neural Networks (VNNs) to seek a latent representation of multi-modal data whose features are jointly captured by a union of subspaces. The so-called self-representation embedding of the latent codes leads to a simplified fusion which is driven by a similarly constructed decoding. The Volterra Filter architecture achieved reduction in parameter complexity is primarily due to controlled non-linearities being introduced by the higher-order convolutions in lieu of generalized activation functions. Experimental results on two different datasets have shown a significant improvement in the clustering performance for VNNs auto-encoder over conventional Convolutional Neural Networks (CNNs) auto-encoder. In addition, we also show that the proposed approach demonstrates a much-improved sample complexity over CNN-based auto-encoder with a robust classification performance.

97 MATHEMATICS AND COMPUTING↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

Generation of Pseudomonas putida KT2440 Strains with Efficient Utilization of Xylose and Galactose via Adaptive Laboratory Evolution

While Pseudomonas putida KT2440 has great potential for biomass-converting processes, its inability to utilize the biomass abundant sugars xylose and galactose has limited its applications. Here, in this study, we utilized Adaptive Laboratory Evolution (ALE) to optimize engineered KT2440 with heterologous expression of xylD encoding xylonate dehydratase from Caulobacter crescentus and galETKM encoding UDP-glucose 4-epimerase, galactose-1-phosphate uridylyltransferase, galactokinase, and galactose-1-epimerase from Escherichia coli K-12 MG1655. Poor starting strain growth (<0.1 h –1 or none) was evolutionarily optimized to rates of up to 0.25 h –1 on xylose and 0.52 h –1 on galactose. Whole-genome sequencing, transcriptomic analysis, and growth screens revealed significant roles of kguT encoding a 2-ketogluconate operon repressor and 2-ketogluconate transporter, and gtsABCD encoding an ATP-binding cassette (ABC) sugar transporting system in xylose and galactose growth conditions, respectively. Finally, we expressed the heterologous indigoidine production pathway in the evolved and unevolved engineered strains and successfully produced 3.2 g/L and 2.2 g/L from 10 g/L of either xylose or galactose in the evolved strains whereas the unevolved strains did not produce any detectable product. Thus, the generated KT2440 strains have the potential for broad application as optimized platform chassis to develop efficient microorganism-based biomass-utilizing bioprocesses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated Symbolic Upscaling: 2. Model Generation for Extended Applicability Regimes

Abstract In this second part of the two paper series, we detail an algorithmic procedure for systematically implementing the generalized closure form strategy presented in Part 1. This strategy extends the applicability of homogenized models with respect to classical homogenization theory, as demonstrated in Part 1 where upscaled models are rigorously derived in moderately reactive physical regimes. After encoding the algorithm into Symbolica, an automated upscaling framework, we upscale two reactive mass transport problems and numerically validate the resulting nonlinear homogenized models by showing the absolute error estimates predicted by homogenization theory are satisfied. In both problems, nontrivial closure forms and closure problems are automatically formulated using the encoded strategy with no human interaction, nor prior knowledge regarding the closure required for the systems. We hope these demonstrations spark further interest in automated analytical frameworks for multiscale modeling, as such capabilities are invaluable for generating rigorous multiscale models of complex phenomena in porous media.

Pietrzyk, Kyle↗

Macroecological distributions of gene variants highlight the functional organization of soil microbial systems

Abstract The recent application of macroecological tools and concepts has made it possible to identify consistent patterns in the distribution of microbial biodiversity, which greatly improved our understanding of the microbial world at large scales. However, the distribution of microbial functions remains largely uncharted from the macroecological point of view. Here, we used macroecological models to examine how the genes encoding the functional capabilities of microorganisms are distributed within and across soil systems. Models built using functional gene array data from 818 soil microbial communities showed that the occupancy-frequency distributions of genes were bimodal in every studied site, and that their rank-abundance distributions were best described by a lognormal model. In addition, the relationships between gene occupancy and abundance were positive in all sites. This allowed us to identify genes with high abundance and ubiquitous distribution (core) and genes with low abundance and limited spatial distribution (satellites), and to show that they encode different sets of microbial traits. Common genes encode microbial traits related to the main biogeochemical cycles (C, N, P and S) while rare genes encode traits related to adaptation to environmental stresses, such as nutrient limitation, resistance to heavy metals and degradation of xenobiotics. Overall, this study characterized for the first time the distribution of microbial functional genes within soil systems, and highlight the interest of macroecological models for understanding the functional organization of microbial systems across spatial scales.

59 BASIC BIOLOGICAL SCIENCES↗

Logical quantum processor based on reconfigurable atom arrays

Suppressing errors is the central challenge for useful quantum computing, requiring quantum error correction (QEC) for large-scale processing. However, the overhead in the realization of error-corrected ‘logical’ qubits, in which information is encoded across many physical qubits for redundancy, poses substantial challenges to large-scale logical quantum computing. Here we report the realization of a programmable quantum processor based on encoded logical qubits operating with up to 280 physical qubits. Using logical-level control and a zoned architecture in reconfigurable neutral-atom arrays, our system combines high two-qubit gate fidelities, arbitrary connectivity, as well as fully programmable single-qubit rotations and mid-circuit readout. Operating this logical processor with various types of encoding, we demonstrate improvement of a two-qubit logic gate by scaling surface-code distance from d = 3 to d = 7, preparation of colour-code qubits with break-even fidelities, fault-tolerant creation of logical Greenberger–Horne–Zeilinger (GHZ) states and feedforward entanglement teleportation, as well as operation of 40 colour-code qubits. Finally, using 3D [[8,3,2]] code blocks, we realize computationally complex sampling circuits with up to 48 logical qubits entangled with hypercube connectivity with 228 logical two-qubit gates and 48 logical CCZ gates. We find that this logical encoding substantially improves algorithmic performance with error detection, outperforming physical-qubit fidelities at both cross-entropy benchmarking and quantum simulations of fast scrambling. These results herald the advent of early error-corrected quantum computation and chart a path towards large-scale logical processors.

74 ATOMIC AND MOLECULAR PHYSICS↗

Unsupervised machine learning discovery of structural units and transformation pathways from imaging data

We show that unsupervised machine learning can be used to learn chemical transformation pathways from observational Scanning Transmission Electron Microscopy (STEM) data. To enable this analysis, we assumed the existence of atoms, a discreteness of atomic classes, and the presence of an explicit relationship between the observed STEM contrast and the presence of atomic units. With only these postulates, we developed a machine learning method leveraging a rotationally invariant variational autoencoder (VAE) that can identify the existing molecular fragments observed within a material. The approach encodes the information contained in STEM image sequences using a small number of latent variables, allowing the exploration of chemical transformation pathways by tracing the evolution of atoms in the latent space of the system. The results suggest that atomically resolved STEM data can be used to derive fundamental physical and chemical mechanisms involved, by providing encodings of the observed structures that act as bottom-up equivalents of structural order parameters. The approach also demonstrates the potential of variational (i.e., Bayesian) methods in the physical sciences and will stimulate the development of more sophisticated ways to encode physical constraints in the encoder–decoder architectures and generative physical laws and causal relationships in the latent space of VAEs.

97 MATHEMATICS AND COMPUTING↗

Key Questions for the Quantum Machine Learner to Ask Themselves

Within the last several years quantum machine learning (QML) has begun to mature; however, many open questions remain. Rather than review open questions, in this perspective piece I will discuss my view about how we should approach problems in QML. In particular I will list a series of questions that I think we should ask ourselves when developing quantum algorithms for machine learning. These questions focus on what the definition of quantum ML is, what is the proper quantum analogue of QML algorithms is, how one should compare QML to traditional ML and what fundamental limitations emerge when trying to build QML protocols. As an illustration of this process I also provide information theoretic arguments that show that amplitude encoding can require exponentially more queries to a quantum model to determine membership of a vector in a concept class than classical bit-encodings would require; however, if the correct analogue is chosen then both the quantum and classical complexities become polynomially equivalent. This example underscores the importance of asking ourselves the right questions when developing and benchmarking QML algorithms.

Wiebe, Nathan O.↗

Machine learning-based real-time kinetic profile reconstruction in DIII-D

Abstract Kinetic equilibrium reconstruction plays a vital role in the physical analysis of plasma stability and control in fusion tokamaks. However, the traditional approach is subjective and prone to human biases. To address this, the consistent automatic kinetic equilibrium reconstruction (CAKE) method was introduced, providing objective results. Nonetheless, its offline nature limits its application in real-time plasma control systems (PCSs). To address this limitation, we present RTCAKENN, a machine learning model that approximates 7 CAKE-level output profiles, namely pressure, inverse q , toroidal current density, electron temperature and density, carbon ion impurity temperature and rotation profiles, using real-time available inputs. The deep neural network consists of an encoder layer, where the scalars and interdependent inputs such as plasma boundary coordinates and motional Stark effect data are encoded using multi-layer perceptrons (MLPs), while profile inputs are encoded by 1D convolutional layers. The encoded data is passed through a MLP for latent feature extraction, before being decoded in the decoding layers, which consist of upsampling and convolutional layers. RTCAKENN has been implemented in the DIII-D PCS and our model achieves accuracy comparable to CAKE and surpasses existing real-time alternatives. Through clever dropout training, RTCAKENN exhibits robustness and can operate even in the absence of Thomson scattering data or charge exchange recombination data. It executes in under 8 ms in the real-time environment, enabling future application in real-time control and analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Phosphate Limitation Responses in Marine Green Algae Are Linked to Reprogramming of the tRNA Epitranscriptome and Codon Usage Bias

Marine algae are central to global carbon fixation, and their productivity is dictated largely by resource availability. Reduced nutrient availability is predicted for vast oceanic regions as an outcome of climate change; however, there is much to learn regarding response mechanisms of the tiny picoplankton that thrive in these environments, especially eukaryotic phytoplankton. Here, we investigate responses of the picoeukaryote Micromonas commoda, a green alga found throughout subtropical and tropical oceans. Under shifting phosphate availability scenarios, transcriptomic analyses revealed altered expression of transfer RNA modification enzymes and biased codon usage of transcripts more abundant during phosphate-limiting versus phosphate-replete conditions, consistent with the role of transfer RNA modifications in regulating codon recognition. To associate the observed shift in the expression of the transfer RNA modification enzyme complement with the transfer RNAs encoded by M. commoda, we also determined the transfer RNA repertoire of this alga revealing potential targets of the modification enzymes. Codon usage bias was particularly pronounced in transcripts encoding proteins with direct roles in managing phosphate limitation and photosystem-associated proteins that have ill-characterized putative functions in “light stress.” The observed codon usage bias corresponds to a proposed stress response mechanism in which the interplay between stress-induced changes in transfer RNA modifications and skewed codon usage in certain essential response genes drives preferential translation of the encoded proteins. Collectively, we expose a potential underlying mechanism for achieving growth under enhanced nutrient limitation that extends beyond the catalog of up- or downregulated protein-encoding genes to the cell biological controls that underpin acclimation to changing environmental conditions.

59 BASIC BIOLOGICAL SCIENCES↗