Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Cyanobacteria from marine oxygen-deficient zones encode both form I and form II Rubiscos

Cyanobacteria are highly abundant in the marine photic zone and primary drivers of the conversion of inorganic carbon into biomass. To date, all studied cyanobacterial lineages encode carbon fixation machinery relying upon form I Rubiscos within a CO 2 -concentrating carboxysome. Here, we report that the uncultivated anoxic marine zone (AMZ) IB lineage ofProchlorococcusfrom pelagic oxygen-deficient zones (ODZs) harbors both form I and form II Rubiscos, the latter of which are typically noncarboxysomal and possess biochemical properties tuned toward low-oxygen environments. We demonstrate that these cyanobacterial form II enzymes are functional in vitro and were likely acquired from proteobacteria. Metagenomic analysis reveals that AMZ IB are essentially restricted to ODZs in the Eastern Pacific, suggesting that form II acquisition may confer an advantage under low-O 2 conditions. AMZ IB populations express both forms of Rubisco in situ, with the highest form II expression at depths where oxygen and light are low, possibly as a mechanism to increase the efficiency of photoautotrophy under energy limitation. Our findings expand the diversity of carbon fixation configurations in the microbial world and may have implications for carbon sequestration in natural and engineered systems.

Science & Technology - Other Topics↗

State-dependent motion of a genetically encoded fluorescent biosensor

Genetically encoded biosensors can measure biochemical properties such as small-molecule concentrations with single-cell resolution, even in vivo. Despite their utility, these sensors are “black boxes”: Very little is known about the structures of their low- and high-fluorescence states or what features are required to transition between them. We used LiLac, a lactate biosensor with a quantitative fluorescence-lifetime readout, as a model system to address these questions. X-ray crystal structures and engineered high-affinity metal bridges demonstrate that LiLac exhibits a large interdomain twist motion that pulls the fluorescent protein away from a “sealed,” high-lifetime state in the absence of lactate to a “cracked,” low-lifetime state in its presence. Understanding the structures and dynamics of LiLac will help to think about and engineer other fluorescent biosensors.

Rosen, Paul C. (ORCID:000000017414454X)↗

Insights into the transcriptional regulation of poorly characterized alcohol acetyltransferase-encoding genes (HgAATs) shed light into the production of acetate esters in the wine yeast Hanseniaspora guilliermondii

Abstract Hanseniaspora guilliermondii is a well-recognized producer of acetate esters associated with fruity and floral aromas. The molecular mechanisms underneath this production or the environmental factors modulating it remain unknown. Herein, we found that, unlike Saccharomyces cerevisiae, H. guilliermondii over-produces acetate esters and higher alcohols at low carbon-to-assimilable nitrogen (C:N) ratios, with the highest titers being obtained in the amino acid-enriched medium YPD. The evidences gathered support a model in which the strict preference of H. guilliermondii for amino acids as nitrogen sources results in a channeling of keto-acids obtained after transamination to higher alcohols and acetate esters. This higher production was accompanied by higher expression of the four HgAATs, genes, recently proposed to encode alcohol acetyl transferases. In silico analyses of these HgAat’s reveal that they harbor conserved AATs motifs, albeit radical substitutions were identified that might result in different kinetic properties. Close homologues of HgAat2, HgAat3, and HgAat4 were only found in members of Hanseniaspora genus and phylogenetic reconstruction shows that these constitute a distinct family of Aat’s. These results advance the exploration of H. guilliermondii as a bio-flavoring agent providing important insights to guide future strategies for strain engineering and media manipulation that can enhance production of aromatic volatiles.

Seixas, Isabel↗

Unified Medical Language System resources improve sieve-based generation and Bidirectional Encoder Representations from Transformers (BERT)–based ranking for concept normalization

Concept normalization, the task of linking phrases in text to concepts in an ontology, is useful for many downstream tasks including relation extraction, information retrieval, etc. We present a generate-and-rank concept normalization system based on our participation in the 2019 National NLP Clinical Challenges Shared Task Track 3 Concept Normalization. The shared task provided 13 609 concept mentions drawn from 100 discharge summaries. We first design a sieve-based system that uses Lucene indices over the training data, Unified Medical Language System (UMLS) preferred terms, and UMLS synonyms to generate a list of possible concepts for each mention. We then design a listwise classifier based on the BERT (Bidirectional Encoder Representations from Transformers) neural network to rank the candidate concepts, integrating UMLS semantic types through a regularizer. Our generate-and-rank system was third of 33 in the competition, outperforming the candidate generator alone (81.66% vs 79.44%) and the previous state of the art (76.35%). During postevaluation, the model’s accuracy was increased to 83.56% via improvements to how training data are generated from UMLS and incorporation of our UMLS semantic type regularizer. Analysis of the model shows that prioritizing UMLS preferred terms yields better performance, that the UMLS semantic type regularizer results in qualitatively better concept predictions, and that the model performs well even on concepts not seen during training. Our generate-and-rank framework for UMLS concept normalization integrates key UMLS features like preferred terms and semantic types with a neural network–based ranking model to accurately link phrases in text to UMLS concepts.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A variational encoder–decoder approach to precise spectroscopic age estimation for large Galactic surveys

Constraints on the formation and evolution of the Milky Way Galaxy require multidimensional measurements of kinematics, abundances, and ages for a large population of stars. Ages for luminous giants, which can be seen to large distances, are an essential component of studies of the Milky Way, but they are traditionally very difficult to estimate precisely for a large data set and often require careful analysis on a star-by-star basis in asteroseismology. Because spectra are easier to obtain for large samples, being able to determine precise ages from spectra allows for large age samples to be constructed, but spectroscopic ages are often imprecise and contaminated by abundance correlations. Here we present an application of a variational encoder–decoder on cross-domain astronomical data to solve these issues. The model is trained on pairs of observations from APOGEE and Kepler of the same star in order to reduce the dimensionality of the APOGEE spectra in a latent space while removing abundance information. The low dimensional latent representation of these spectra can then be trained to predict age with just ∼1000 precise seismic ages. We demonstrate that this model produces more precise spectroscopic ages (∼ 22 per cent overall, ∼ 11 per cent for red-clump stars) than previous data-driven spectroscopic ages while being less contaminated by abundance information (in particular, our ages do not depend on [α/M]). We create a public age catalogue for the APOGEE DR17 data set and use it to map the age distribution and the age-[Fe/H]-[α/M] distribution across the radial range of the Galactic disc.

79 ASTRONOMY AND ASTROPHYSICS↗

Random circuit block-encoded matrix and a proposal of quantum LINPACK benchmark

The LINPACK benchmark reports the performance of a computer for solving a system of linear equations with dense random matrices. Although this task was not designed with a real application directly in mind, the LINPACK benchmark has been used to define the list of TOP500 supercomputers since the debut of the list in 1993. We propose that a similar benchmark, called the quantum LINPACK benchmark, could be used to measure the whole machine performance of quantum computers. The success of the quantum LINPACK benchmark should be viewed as the minimal requirement for a quantum computer to perform a useful task of solving linear algebra problems, such as linear systems of equations. We propose an input model called the Random Circuit Block-Encoded Matrix (RACBEM), which is a proper generalization of a dense random matrix in the quantum setting. The RACBEM model is efficient to be implemented on a quantum computer and can be designed to optimally adapt to any given quantum architecture, with relying on a black-box quantum compiler. Besides solving linear systems, the RACBEM model can be used to perform a variety of linear algebra tasks relevant to many physical applications, such as computing spectral measures, time series generated by a Hamiltonian simulation, and thermal averages of the energy. We implement these linear algebra operations on IBM Q quantum devices as well as quantum virtual machines, and demonstrate their performance in solving scientific computing problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Exact block encoding of imaginary time evolution with universal quantum neural networks

We develop a constructive approach to generate quantum neural networks capable of representing the exact thermal states of all many-body qubit Hamiltonians. The Trotter expansion of the imaginary time propagator is implemented through an exact block encoding by means of a unitary, restricted Boltzmann machine architecture. Marginalization over the hidden-layer neurons (auxiliary qubits) creates the nonunitary action on the visible layer. Then, we introduce a unitary deep Boltzmann machine architecture in which the hidden-layer qubits are allowed to couple laterally to other hidden qubits. We prove that this wave-function is closed under the action of the imaginary time propagator and, more generally, can represent the action of a universal set of quantum gate operations. We provide analytic expressions for the coefficients for both architectures, thus enabling exact network representations of thermal states without stochastic optimization of the network parameters. In the limit of large imaginary time, the yields the ground state of the system. The number of qubits grows linearly with the number of interactions and total imaginary time for a fixed interaction order. Both networks can be readily implemented on quantum hardware via midcircuit measurements of auxiliary qubits. If only one auxiliary qubit is measured and reset, the circuit depth scales linearly with imaginary time and number of interactions, while the width is constant. Alternatively, one can employ a number of auxiliary qubits linearly proportional to the number of interactions, and circuit depth grows linearly with imaginary time only. Every midcircuit measurement has a postselection success probability, and the overall success probability is equal to the product of the probabilities of the midcircuit measurements.

97 MATHEMATICS AND COMPUTING↗

Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal Runaway

Lithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1.

Tabassum, Anika [ORNL] (ORCID:0000000254600955)↗

Hardware-Based Randomized Encoding for Sensor Authentication in Power Grid SCADA Systems

Supervisory Control and Data Acquisition (SCADA) systems are utilized extensively in critical power grid infrastructures. Modern SCADA systems have been proven to be susceptible to cyber-security attacks and require improved security primitives in order to prevent unwanted influence from an adversarial party. One section of weakness in the SCADA system is the integrity of field level sensors providing essential data for control decisions at a master station. In this paper we propose a lightweight hardware scheme providing inferred authentication for SCADA sensors by combining an analog to digital converter and a permutation generator as a single integrated circuit. Through this method we encode critical sensor data at the time of sensing, so that unencoded data is never stored in memory, increasing the difficulty of software attacks. We show through experimentation how our design stops both software and hardware false data injection attacks occurring at the field level of SCADA systems.

42 ENGINEERING↗

PHOTO‐SENSITIVE LEAF ROLLING 1 encodes a polygalacturonase that modifies cell wall structure and drought tolerance in rice

Summary The biosynthesis and modification of cell wall composition and structure are controlled by hundreds of enzymes and have a direct consequence on plant growth and development. However, the majority of these enzymes has not been functionally characterised. Rice mutants with leaf‐rolling phenotypes were screened in a field. Phenotypic analysis under controlled conditions was performed for the selected mutant and the relevant gene was identified by map‐based cloning. Cell wall composition was analysed by glycome profiling assay. We identified a photo‐sensitive leaf rolling 1 ( psl1 ) mutant with ‘napping’ (midday depression of photosynthesis) phenotype and reduced growth. The PSL1 gene encodes a cell wall‐localised polygalacturonase (PG), a pectin‐degrading enzyme. psl1 with a 260‐bp deletion in its gene displayed leaf rolling in response to high light intensity and/or low humidity. Biochemical assays revealed PG activity of recombinant PSL1 protein. Significant modifications to cell wall composition in the psl1 mutant compared with the wild‐type plants were identified. Such modifications enhanced drought tolerance of the mutant plants by reducing water loss under osmotic stress and drought conditions. Taken together, PSL1 functions as a PG that modifies cell wall biosynthesis, plant development and drought tolerance in rice.

Zhang, Guangheng↗

Amino acid substrate specificities and tissue expression profiles of the nine CYP79A encoding genes in Sorghum bicolor

Cytochrome P450s of the CYP79 family catalyze two N-hydroxylation reactions, converting a selected number of amino acids into the corresponding oximes. The sorghum genome (Sorghum bicolor) harbours nine CYP79A encoding genes, and here sequence comparisons of the CYP79As along with their substrate recognition sites (SRSs) are provided. The substrate specificity of previously uncharacterized CYP79As was investigated by transient expression in Nicotiana benthamiana and subsequent transformation of the oximes formed into the corresponding stable oxime glucosides catalyzed by endogenous UDPG-glucosyltransferases (UGTs). CYP79A61 uses phenylalanine as a substrate, whereas CYP79A91, CYP79A93, and CYP79A95 use valine and isoleucine as substrates, with CYP79A93 showing the ability also to use phenylalanine. CYP79A94 uses isoleucine as a substrate. Analysis of 249 sorghum transcriptomes from two different sorghum cultivars showed the expression levels and tissue-specific expression of the CYP79As. CYP79A1 is the committed gene in dhurrin formation and was the highest expressed gene in most tissues/organs. CYP79A61 was primarily expressed in fully developed leaf blades and leaf sheaths. CYP79A91 and CYP79A92 were expressed mainly in roots >200 cm below ground, while CYP79A93 and CYP79A94 were most highly expressed in the leaf collar and leaf sheath, respectively. Here, the possible signalling effects of the oximes and their metabolites produced in different sorghum tissues are discussed.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-Guided, Physics-Informed, and Physics-Encoded Neural Networks and Operators in Scientific Computing: Fluid and Solid Mechanics

Abstract Advancements in computing power have recently made it possible to utilize machine learning and deep learning to push scientific computing forward in a range of disciplines, such as fluid mechanics, solid mechanics, materials science, etc. The incorporation of neural networks is particularly crucial in this hybridization process. Due to their intrinsic architecture, conventional neural networks cannot be successfully trained and scoped when data are sparse, which is the case in many scientific and engineering domains. Nonetheless, neural networks provide a solid foundation to respect physics-driven or knowledge-based constraints during training. Generally speaking, there are three distinct neural network frameworks to enforce the underlying physics: (i) physics-guided neural networks (PgNNs), (ii) physics-informed neural networks (PiNNs), and (iii) physics-encoded neural networks (PeNNs). These methods provide distinct advantages for accelerating the numerical modeling of complex multiscale multiphysics phenomena. In addition, the recent developments in neural operators (NOs) add another dimension to these new simulation paradigms, especially when the real-time prediction of complex multiphysics systems is required. All these models also come with their own unique drawbacks and limitations that call for further fundamental research. This study aims to present a review of the four neural network frameworks (i.e., PgNNs, PiNNs, PeNNs, and NOs) used in scientific computing research. The state-of-the-art architectures and their applications are reviewed, limitations are discussed, and future research opportunities are presented in terms of improving algorithms, considering causalities, expanding applications, and coupling scientific and deep learning solvers.

Computer Science↗

A bistable and reconfigurable molecular system with encodable bonds

Molecular systems with ability to controllably transform between different conformations play pivotal roles in regulating biochemical functions. Here, we report the design of a bistable DNA origami four-way junction (DOJ) molecular system that adopts two distinct stable conformations with controllable reconfigurability by using conformation-controlled base stacking. Exquisite control over DOJ’s conformation and transformation is realized by programming the stacking bonds (quasi–blunt-ends) within the junction to induce prescribed coaxial stacking of neighboring junction arms. A specific DOJ conformation may be achieved by encoding the stacking bonds with binary stacking sequences based on thermodynamic calculations. Dynamic transformations of DOJ between various conformations are achieved by using specific environmental and molecular stimulations to reprogram the stacking codes. This work provides a useful platform for constructing self-assembled DNA nanostructures and nanomachines and insights for future design of artificial molecular systems with increasing complexity and reconfigurability.

60 APPLIED LIFE SCIENCES↗

Structure of a Vaccine-Induced, Germline-Encoded Human Antibody Defines a Neutralizing Epitope on the SARS-CoV-2 Spike N-Terminal Domain

Structural characterization of infection- and vaccination-elicited antibodies in complex with antigen provides insight into the evolutionary arms race between the host and the pathogen and informs rational vaccine immunogen design. We isolated a germ line-encoded monoclonal antibody (mAb) from plasmablasts activated upon mRNA vaccination against severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and determined its structure in complex with the spike glycoprotein by electron cryomicroscopy (cryo-EM). We show that the mAb engages a previously uncharacterized neutralizing epitope on the spike N-terminal domain (NTD). The high-resolution structure reveals details of the intermolecular interactions and shows that the mAb inserts its heavy complementarity-determining region 3 (HCDR3) loop into a hydrophobic NTD cavity previously shown to bind a heme metabolite, biliverdin. We demonstrate direct competition with biliverdin and that, because of the conserved nature of the epitope, the mAb maintains binding to viral variants B.1.1.7 (alpha), B.1.351 (beta), B.1.617.2 (delta), and B.1.1.529 (omicron). Our study describes a novel conserved epitope on the NTD that is readily targeted by vaccine-induced antibody responses.

59 BASIC BIOLOGICAL SCIENCES↗

Tencoder: tensor-product encoder-decoder architecture for predicting solutions of PDEs with variable boundary data

It is widely hoped that artificial intelligence will boost data-driven surrogate models in science and engineering. However, fundamental spatial aspects of AI surrogate models remain under-studied. We investigate the ability of neural-network surrogate models to predict solutions to PDEs under variable boundary values. We do not wish to retrain the model when the boundary values change but to make them inputs to the model and infer the solution of the PDE under those boundary conditions. Such a capability is essential to making AI-based surrogate models practically useful. While simple feedforward networks are used for one-dimensional (1D) Poisson equation, an encoder-decoder architecture with a tensor-product layer is developed for the two-dimensional Poisson equation posed on a rectangular domain. We show that it is indeed possible to infer solutions to PDEs from variable boundary data using neural networks in this relatively simple setting, and point to future directions.

Kashi, Aditya↗

HD-Bind: Encoding of Molecular Structure with Low Precision, Hyperdimensional Binary Representations

Publicly available collections of drug-like molecules have grown to comprise tens of billions of compounds due to advances in combinatorial chemistry. Traditional methods for identifying "hit" molecules from a large collection of potential drug-like candidates have relied on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have the major drawback that they require exceptional computing capabilities for even relatively small collections of molecules. Hyperdimensional Computing (HDC) is a recently-proposed learning paradigm that represents data with high-dimension binary vectors; this allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas. We consider existing HDC approaches for molecular property classification and introduce two novel encodings of a commonly-used molecular representation, the extended connectivity fingerprint (ECFP). We show that HDC-based inference methods are as much as 91 times more efficient than traditional machine learning methods, and achieve an acceleration of nearly nine orders of magnitude compared to molecular docking. Our results show that HDC accelerated methods retain competitive accuracy on a number of well-studied tasks such as molecular property predictions using the MoleculeNet dataset, and bind/no-bind activity classification using the DUD-E and LIT-PCBA datasets. Our work thus motivates further investigation into molecular representation learning to develop ultraefficient pre-screening tools.

Jones, William↗

Enhancing Short-Range Weather Forecasts through Temporal Variation Encoding: A Multiperiod Embedding Approach

Machine learning (ML) techniques have emerged as promising approaches to improve regional weather forecast accuracy and reliability through data-driven methods. We propose a novel ML-based weather forecasting model, the Multiperiod Embed Net (MPENet). A key distinguishing feature of MPENet is its explicit utilization of the inherent cyclic nature in weather dynamics, unlike the autoregressive strategies commonly used in other ML weather forecasting approaches. Critical cyclic structures are identified via Fourier analyses of dynamic time series. Cyclicity in the convolutional representation is achieved by transforming one-dimensional time series of meteorological variables into two-dimensional tensors based on identified periods. This approach enables the model to leverage intrinsic weather patterns, enhancing regional forecast performance. To demonstrate the effectiveness of MPENet, we conduct a comparative analysis with Nvidia’s FourCastNet. Both models are trained on High-Resolution Rapid Refresh (HRRR) data from 2015 to 2022, over a 192 km × 192 km region in Tennessee. The comparisons are performed locally at two specific locations known to have different weather dynamics due to orographic effects: Crossville, on the relatively flat Cumberland Plateau with fewer topographic airflow disruptions, and Oak Ridge, in the ridge-and-valley region, where airflow is heavily influenced by surrounding valleys and mountains. Our results indicate that FourCastNet achieves strong accuracy at very short lead times, while MPENet maintains competitive skill and shows advantages in capturing temporal evolution over longer periods. Cross-correlation analyses of MPENet and FourCastNet predictions with the HRRR data suggest that encoding critical cyclicity into the network architecture leads to improvements in the forecasting skill.

Artificial intelligence↗

Encoding the complete electric field of an ultraviolet ultrashort laser pulse in a near-infrared nonlinear-optical signal

We introduce a variation on the cross-correlation frequency-resolved optical gating (XFROG) technique that uses a near-infrared (NIR) nonlinear-optical signal to characterize pulses in the ultraviolet (UV). Using a transient-grating XFROG beam geometry, we create a grating using two copies of the unknown UV pulse and diffract a NIR reference pulse from it. We show that, by varying the delay between the UV pulses creating the grating, the UV pulse intensity-and-phase information can be encoded into a NIR signal. We also implemented a modified generalized-projections phase-retrieval algorithm for retrieving the UV pulses from these spectrograms. We performed proof-of-principle measurements of chirped pulses and double pulses, all at 400 nm. This approach should be extendable deeper into the UV and potentially even into the extreme UV or x-ray range.

47 OTHER INSTRUMENTATION↗