Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data structures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Consistent performance of large language models in rare disease diagnosis across ten languages and 4917 cases

Background Large language models (LLMs) are increasingly used medicine for diverse applications including differential diagnostic support. The training data used to create LLMs such as the Generative Pretrained Transformer (GPT) predominantly consist of English-language texts, but LLMs could be used across the globe to support diagnostics if language barriers could be overcome. Initial pilot studies on the utility of LLMs for differential diagnosis in languages other than English have shown promise, but a large-scale assessment on the relative performance of these models in a variety of European and non-European languages on a comprehensive corpus of challenging rare-disease cases is lacking. Methods We created 4917 clinical vignettes using structured data captured with Human Phenotype Ontology (HPO) terms with the Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema. These clinical vignettes span a total of 360 distinct genetic diseases with 2525 associated phenotypic features. We used translations of the Human Phenotype Ontology together with language-specific templates to generate prompts in English, Chinese, Czech, Dutch, French, German, Italian, Japanese, Spanish, and Turkish. We applied GPT-4o, version gpt-4o-2024-08-06, and the medically fine-tuned Meditron3-70B to the task of delivering a ranked differential diagnosis using a zero-shot prompt. An ontology-based approach with the Mondo disease ontology was used to map synonyms and to map disease subtypes to clinical diagnoses in order to automate evaluation of LLM responses. Findings For English, GPT-4o placed the correct diagnosis at the first rank 19.9% and within the top-3 ranks 27.0% of the time. In comparison, for the nine non-English languages tested here the correct diagnosis was placed at rank 1 between 16.9% and 20.6%, within top-3 between 25.4% and 28.6% of cases. The Meditron3 model placed the correct diagnosis within the first 3 ranks for 20.9% of cases in English and between 19.9% and 24.0% for the other nine languages. Interpretation The differential diagnostic performance of LLMs across a comprehensive corpus of rare-disease cases was largely consistent across the ten languages tested. This suggests that the utility of LLMs in clinical settings may extend to non-English clinical settings.

Artificial intelligence

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING

ED-cPSD: Fast Phase-Size Distribution via Sequential Erosion-Dilation

The Erosion-Dilation continuous Phase-Size Distribution, ED-cPSD, is an application for calculating continuous pore and particle-size distribution from digital reconstructions and/or image-based structural data. It is based on the erosion-dilation continuous phase-size distribution method. A continuous size distribution is a measure of the probability density of finding a particle or pore of a certain size. These distributions are of interest in any field of study involving porous media, including but not limited to electrochemistry, petroleum engineering, geology, and food science. The algorithm behind the software provides a computationally efficient way to calculate phase-size distributions for large domains. For a 3D battery electrode reconstruction with 1.3 x 10 8 voxels, the particle size distribution is derived in under 2 min on a desktop, while also retaining flexibility and computational efficiency for HPC-scale multi-threading. The software can handle structures with over 10 9 voxels. The algorithm is roughly 280 times faster than a previous version on the same task.

Characterization

srlife : A software tool for estimating the life of high temperature concentrating solar receivers. Part II – Ceramic receivers

As Concentrating Solar Power (CSP) technologies aim for higher operating temperatures to enhance efficiency and meet industrial process heat demands, high-temperature metallic materials, including nickel-based superalloys, face challenges in maintaining structural integrity. Advanced ceramics offer a promising alternative due to their superior high-temperature strength. However, accurately assessing the performance of ceramic components requires a fundamentally different approach from that used for metallic components. This Part II of a two-part paper describes the integration of ceramic statistical failure models within srlife – an open-source tool for predicting the life of high-temperature CSP receivers. These models account for the inherent variability in ceramic strength, as well as the effects of subcritical crack growth (SCG) under high temperature cyclic loads. Here, the paper includes an example problem that demonstrate the process of evaluating ceramic receivers using srlife. Part I details the life estimation process for metallic receivers (i.e. creep-fatigue life) along with input and output data structure, thermohydraulic analysis, and structural analysis. The complete tool is available as open-source software at https://github.com/srlife-project/srlife and can be installed via the PyPi package manager (https://pypi.org). By supporting both ceramic and metallic receiver analyses, srlife facilitates fair comparisons between competing metallic and ceramic designs, enabling accurate evaluations of plant efficiency and the economic benefits of ceramic solar receivers and other components.

High temperature ceramic receivers

Discrete global grid system-based flow routing datasets in the Amazon and Yukon basins

Abstract. Discrete global grid systems (DGGS) are emerging spatial data structures widely used to organize geospatial datasets across scales. While DGGS have found applications in various scientific disciplines, including atmospheric science and ecology, their integration into physically based hydrological models and Earth system models (ESMs) has been hindered by the lack of flow routing datasets based on DGGS. In response to this gap, this study pioneers the development of new flow routing datasets using icosahedral Snyder equal-area (ISEA) DGGS and a novel mesh-independent flow direction model. We present flow routing datasets for two large basins, the tropical Amazon River basin and the Arctic Yukon River basin. These datasets (1) facilitate the adoption of DGGS for hydrological models and (2) provide flow routing inputs for evaluation of DGGS-based flow routing in the Amazon and Yukon river basins. The data are available at https://doi.org/10.5281/zenodo.8377765 (Liao, 2023).

54 ENVIRONMENTAL SCIENCES

Understanding Strain and Failure of a Knot in Polyethylene Using Molecular Dynamics with Machine-Learned Potentials

A neural network potential (NNP) has been developed by fitting to ab initio electronic structure data on hydrocarbons and is used to study failure of linear and knotted polyethylene (PE) chains. A linear PE chain must be highly strained before breaking as the stress is equally distributed across the chain. In contrast, the stress in a PE chain with a 31 or overhand knot, accumulates at the knot’s entrance/exit. We find the strain energy is greatest when the bond length and angle are strained simultaneously, and that the knot weakens the chain by increasing the variance of the C–C–C angle, thereby allowing rupture at lower bond strains. Here, we extend our analysis to both 51 and 52 knots and find that both break at the entrance/exit of a loop. Notably, molecular scale PE knots exhibit many of the same characteristics as knots in a macroscopic rope, with stick–slip phenomena upon tightening and similar points of failure.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Comprehensive review of 2 β decay half-lives

Here, the double-beta (2 β )-decay is the rarest nuclear physics process, and its experimental half-lives (T 1/2 ) exceed the age of the Universe from nine to fourteen orders of magnitude. Double-beta decay was observed, and its half-life was measured in 14 parent nuclei using direct, radiochemical, and geochemical methods. The decay observables are analyzed using the Evaluated Nuclear Structure Data File (ENSDF) procedures, and the recommended T 1/2 were deduced. Using the calculated values of phase factors, the effective nuclear matrix elements were extracted and compared with available data. Thousands of theoretical and experimental works have been dedicated to these topics in the last 85 years, and we present two data sets of recommended values to encapsulate the results.

2β-decay

Visualization of the surface distribution of photocatalytic activity at the anatase/rutile TiO2 interface using X-ray photoelectron emission microscopy

We applied X-ray photoelectron emission microscopy (XPEEM) to visualize the surface distribution of photocatalytic activity at the anatase/rutile (A/R) interface of titanium dioxide. Acetic acid was adsorbed on the surface, and its decomposition and desorption reactions under ultraviolet light irradiation were monitored by observing C 1s spectra, enabling direct observation of the spatial distribution of photocatalytic activity. Complementary crystal structure information was obtained using spatially resolved low-energy electron diffraction, low-energy electron microscopy, X-ray absorption spectroscopy, and micro-focused Raman spectroscopy. Comparison with these structural data allowed for the evaluation of the photocatalytic activity as a function of the distance from the A/R interface. The activity is enhanced in the vicinity of the interface and gradually decreases toward both the anatase and rutile regions. These results demonstrate that XPEEM is an effective technique for spatially resolved mapping of photocatalytic activity, which has\\r\\npreviously been inaccessible using conventional characterization techniques.

25 ENERGY STORAGE

Integrated multi-omic characterizations of the synapse reveal RNA processing factors and ubiquitin ligases associated with neurodevelopmental disorders

The molecular composition of the excitatory synapse is incompletely defined due to its dynamic nature across developmental stages and neuronal populations. To address this gap, we apply proteomic mass spectrometry to characterize the synapse in multiple biological models including the fetal human brain and hiPSC-derived neurons. To prioritize the identified proteins, we develop an orthogonal multi-omic screen of genomic, transcriptomic, interactomic, and structural data. This data-driven framework identifies proteins with key molecular features intrinsic to the synapse, including characteristic patterns of biophysical interactions and cross-tissue expression. The multi-omic analysis captures synaptic proteins across developmental stages and experimental systems, including 493 synaptic candidates supported by proteomics. We further investigate three such proteins that are associated with neurodevelopmental disorders – the CUL3 E3 ubiquitin ligase, the DDX3X and YBX1 nucleic-acid binding proteins – by mapping their networks of physically interacting synapse proteins or transcripts. Our study demonstrates the potential of an integrated multi-omic approach to systematically and more comprehensively resolve the synaptic architecture.

59 BASIC BIOLOGICAL SCIENCES

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben

A High-Performance Discrete-Element Framework for Simulating Flow and Jamming of Moisture Bearing Biomass Feedstocks

We developed and verified a high-performance open-source discrete element method (DEM) solver with simultaneously-supported feedstock-specific interaction models, including bonded-sphere, liquid bridge, cohesion, and non-linear contact models. Our solver uses parallel data structures on hybrid central and graphics processing unit (CPU/GPU) architectures, with favorable strong scaling performance observed for large problem sizes comprised of (100 M particles), and 4X single-node GPU speedup. The particles for corn stover feedstock were conceptualized and calibrated based on experimental measurements and results. Sensitivity analyses demonstrate that the mass flow rate from a wedge hopper is governed primarily by moisture content, friction coefficient, and cohesion energy density. The model is used to reproduce experimentally observed hopper jamming results, highlighting that the experimental no-flow trends can only be achieved by using non-spherical particles, liquid bridge and cohesion models, highlighting the importance of using concurrent feedstock specialized models for the effective representation of biomass material handling problems.

bioenergy

A Spectroscopic and Computational Evaluation of Uranyl Oxo Engagement with Transition Metal Cations

Here, we report the synthesis and characterization of five novel Cd 2+ /UO 2 2+ heterometallic complexes that feature Cd-oxo distances ranging from 78 to 171% of the sum of the van der Waals radii for these atoms. This work marks an extension of our previously reported Pb 2+ /UO 2 2+ and Ag + /UO 2 2+ complexes, yet with much more pronounced structural and spectroscopic effects resulting from Cd-oxo interactions. We observe a major shift in the U═O symmetric stretch and significant uranyl bond length asymmetry. The ρbcp values calculated using Quantum Theory of Atoms in Molecules (QTAIM) support the asymmetry displayed in the structural data and indicate a decrease in covalent character in U═O bonds with close Cd-oxo contacts, more so than in related compounds containing Pb 2+ and Ag + . Second-order perturbation theory (SOPT) analysis reveals that O sp x → Cd s is the most significant orbital overlap and U═O bonding and antibonding orbitals also contribute to the interaction (U═O σ/π → Cd d and Cd s → U═O σ/π*). The overall stabilization energies for these interactions were lower than those in previously reported Pb 2+ cations, yet larger than related Ag + compounds. Analysis of the equatorial coordination sphere of the Cd 2+ /UO 2 2+ compounds (along with Pb 2+ /UO 2 2+ complexes) reveals that 7-coordinate uranium favors closer, stronger M n+ -oxo contacts. These results indicate that U═O bond strength tuning is possible with judicious choice of metal cations for oxo interactions and equatorial ligand coordination.

cations

Functional and Structural Studies on the Esperamicin Thioesterase and Progress toward Understanding Enediyne Core Biosynthesis

Enediynes are among the most potent antitumor and antibacterial natural products. Studies on their biosynthetic pathways have identified a shared, linear polyene precursor generated from an iterative type I polyketide synthase (PKSE) as the source of the enediyne warhead. A key step is the release of this polyene from the PKSE by a discrete thioesterase (TE). Here, in this study, we used X-ray crystallography, site-directed mutagenesis, and heterologous coexpression of PKSEs and TEs to elucidate how enediyne TEs mediate the production of the polyene. We solved the structure of wild-type EspE7 from esperamicin producer Actinomodura verrucosospora. The substrate binding pocket was also defined upon serendipitous cocrystallization of an EspE7 mutant with a fatty acyl-CoA ligand. Structural data and in vitro activity assays with EspE7 mutants provide strong evidence that Glu68 in EspE7 and the analogous Glu residue in other enediyne TEs functions as a key catalytic residue, thus supporting a hydrolysis mechanism for enediyne TEs that aligns with that of Pseudomonas sp. 4-HB-CoA TE. Furthermore, combinations of 9- and 10-membered enediyne PKSEs and TEs produced 1,3,5,7,9,11,13-pentadecaheptaene (1) as the major product. Thus, the data further support previous conclusions that 1 serves as the sole precursor for the biosynthesis of all enediyne cores.

Enediynes

Kinetics and Mechanism of the Singlet Oxygen Atom Reaction with Dimethyl Ether

Here, we combine in situ laser spectroscopy, quantum chemistry, and kinetic calculations to study the reaction of a singlet oxygen atom with dimethyl ether. Infrared laser absorption spectroscopy and Faraday rotation spectroscopy are used for the detection and quantification of the reaction products OH, H 2 O, HO 2 , and CH 2 O on submillisecond time scales. Fitting temporal profiles of products with simulations using an in-house reaction mechanism allows product branching to be quantified at 30, 60, and 150 Torr. The experimentally determined product branching agrees well with master equation calculations based on electronic structure data and transition state theory. The calculations demonstrate that the dimethyl peroxide (CH 3 OOCH 3 ) generated via O-insertion into the C–O bond undergoes subsequent dissociation to CH 3 O + CH 3 O through energetically favored reactions without an intrinsic barrier. This O-insertion mechanism can be important for understanding the fate of biofuels leaking into the atmosphere and for plasma-based biofuel processing technologies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Stability Frontiers in the AM 6 X 6 Kagome Metals: The Ln Nb 6 Sn 6 ( Ln :Ce–Lu,Y) Family and Density-Wave Transition in LuNb 6 Sn 6

The kagome motif is a versatile platform for condensed matter physics, hosting rich interactions between magnetic, electronic, and structural degrees of freedom. In recent years, the discovery of a charge density wave (CDW) in the AV 3 Sb 5 superconductors and structurally-derived bond density waves (BDW) in FeGe and ScV 6 Sn 6 have stoked the search for new kagome platforms broadly exhibiting density wave (DW) transitions. Here, in this work, we evaluate the known AM 6 X 6 chemistries and construct a stability diagram that summarizes the structural relationships among the >125 member family. Subsequently, we introduce our discovery of the broader LnNb 6 Sn 6 (Ln:Ce–Nd,Sm,Gd–Tm,Lu,Y) family of kagome metals and an analogous DW transition in LuNb 6 Sn 6 . Our X-ray scattering measurements clearly indicate a (1/3, 1/3, 1/3) ordering wave vector (√$\bar{3}$ x √$\bar{3}$ x $3$ superlattice) and diffuse scattering on half-integer L-planes. Our analysis of the structural data supports the “rattling mode” DW model proposed for ScV 6 Sn 6 and paints a detailed picture of the steric interactions between the rare-earth filler element and the host Nb–Sn kagome scaffolding. We also provide a broad survey of the magnetic properties within the HfFe 6 Ge 6 -type LnNb 6 Sn 6 members, revealing a number of complex antiferromagnetic and metamagnetic transitions throughout the family. This work integrates our new LnNb 6 Sn 6 series of compounds into the broader AM 6 X 6 family, providing new material platforms and forging a new route forward at the frontier of kagome metal research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Structural basis for C-degron selectivity across KLHDCX family E3 ubiquitin ligases

Abstract Specificity of the ubiquitin-proteasome system depends on E3 ligase-substrate interactions. Many such pairings depend on E3 ligases binding to peptide-like sequences - termed N- or C-degrons - at the termini of substrates. However, our knowledge of structural features distinguishing closely related C-degron substrate-E3 pairings is limited. Here, by systematically comparing ubiquitylation activities towards a suite of common model substrates, and defining interactions by biochemistry, crystallography, and cryo-EM, we reveal principles of C-degron recognition across the KLHDCX family of Cullin-RING ligases (CRLs). First, a motif common across these E3 ligases anchors a substrate’s C-terminus. However, distinct locations of this C-terminus anchor motif in different blades of the KLHDC2, KLHDC3, and KLHDC10 β-propellers establishes distinct relative positioning and molecular environments for substrate C-termini. Second, our structural data show KLHDC3 has a pre-formed pocket establishing preference for an Arg or Gln preceding a C-terminal Gly, whereas conformational malleability contributes to KLHDC10’s recognition of varying features adjacent to substrate C-termini. Finally, additional non-consensus interactions, mediated by C-degron binding grooves and/or by distal propeller surfaces and substrate globular domains, can substantially impact substrate binding and ubiquitylatability. Overall, the data reveal combinatorial mechanisms determining specificity and plasticity of substrate recognition by KLDCX-family C-degron E3 ligases.

Science & Technology - Other Topics

Molecular model of TFIIH recruitment to the transcription-coupled repair machinery

Transcription-coupled repair (TCR) is a vital nucleotide excision repair sub-pathway that removes DNA lesions from actively transcribed DNA strands. Binding of CSB to lesion-stalled RNA Polymerase II (Pol II) initiates TCR by triggering the recruitment of downstream repair factors. Yet it remains unknown how transcription factor IIH (TFIIH) is recruited to the intact TCR complex. Combining existing structural data with AlphaFold predictions, we build an integrative model of the initial TFIIH-bound TCR complex. We show how TFIIH can be first recruited in an open repair-inhibited conformation, which requires subsequent CAK module removal and conformational closure to process damaged DNA. In our model, CSB, CSA, UVSSA, elongation factor 1 (ELOF1), and specific Pol II and UVSSA-bound ubiquitin moieties come together to provide interaction interfaces needed for TFIIH recruitment. STK19 acts as a linchpin of the assembly, orienting the incoming TFIIH and bridging Pol II to core TCR factors and DNA. Molecular simulations of the TCR-associated CRL4CSA ubiquitin ligase complex unveil the interplay of segmental DDB1 flexibility, continuous Cullin4A flexibility, and the key role of ELOF1 for Pol II ubiquitination that enables TCR. Collectively, these findings elucidate the coordinated assembly of repair proteins in early TCR.

Paul, Tanmoy

Efficient generation of grids and traversal graphs in compositional spaces towards exploration and path planning

Abstract Diverse disciplines across science and engineering deal with problems related to compositions, which exist in non-Euclidean simplex spaces, rendering many standard tools inaccurate or inefficient. This work explores such spaces conceptually in the context of materials discovery, quantifies their computational feasibility, and implements several essential methods specific to simplex spaces through a new high-performance open-source library . Most significantly, we derive and implement an algorithm for constructing a novel n-dimensional simplex graph data structure, containing all discretized compositions and possible neighbor-to-neighbor transitions. Critically, no distance or neighborhood calculations are performed, instead leveraging pure combinatorics and order in procedurally generated simplex grids, keeping the algorithm $${\mathcal{O}}(N)$$ O ( N ) , with minimal memory, enabling rapid construction of graphs with billions of transitions in seconds. Additionally, we demonstrate how such graph representations can be combined to homogeneously express complex path-planning problems, while facilitating efficient deployment of existing high-performance gradient descent, graph traversal, and other optimization algorithms.

Krajewski, Adam M. (ORCID:0000000222660099)