Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Human Error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Patient mutations in human ATP:cob(I)alamin adenosyltransferase differentially affect its catalytic versus chaperone functions

Human ATP:cob(I)alamin adenosyltransferase (ATR) is a mitochondrial enzyme that catalyzes an adenosyl transfer to cob(I)alamin, synthesizing 5'-deoxyadenosylcobalamin (AdoCbl) or coenzyme B 12 . ATR is also a chaperone that escorts AdoCbl, transferring it to methylmalonyl-CoA mutase, which is important in propionate metabolism. Mutations in ATR lead to methylmalonic aciduria type B, an inborn error of B12 metabolism. Our previous studies have furnished insights into how ATR protein dynamics influence redox-linked cobalt coordination chemistry, controlling its catalytic versus chaperone functions. In this study, we have characterized three patient mutations at two conserved active site residues in human ATR, R190C/H, and E193K and obtained crystal structures of R190C and E193K variants, which display only subtle structural changes. All three mutations were found to weaken affinities for the cob(II)alamin substrate and the AdoCbl product and increase K M(ATP) . 31 P NMR studies show that binding of the triphosphate product, formed during the adenosylation reaction, is also weakened. However, although the k cat of this reaction is significantly diminished for the R190C/H mutants, it is comparable with the WT enzyme for the E193K variant, revealing the catalytic importance of Arg-190. Furthermore, although the E193K mutation selectively impairs the chaperone function by promoting product release into solution, its catalytic function might be unaffected at physiological ATP concentrations. In contrast, the R190C/H mutations affect both the catalytic and chaperoning activities of ATR. Because the E193K mutation spares the catalytic activity of ATR, our data suggest that the patients carrying this mutation are more likely to be responsive to cobalamin therapy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Eddy covariance towers as sentinels of abnormal radioactive material releases

Ensuring accurate detection and attribution of abnormal releases of radioactive material is critical for protecting human health and safety. Most commonly, such detection is accomplished via active monitoring approaches involving the collection of physical samples. Further, this is labor intensive and limits the temporal and spatial resolution of any detected events to a relatively coarse level. As an alternative first step towards passive monitoring, we developed an approach using eddy flux tower data records to identify signals from a known abnormal release and quantify the extent to which that signal also occurs at other times in the data record. Through two case studies, one of which targeted the Fukushima nuclear disaster and the other targeting an abnormal release event at a radioisotope production facility in Fleurus, Belgium, we tested our approach and identified several potential heretofore unidentified abnormal events that were consistent with atmospheric circulation patterns and/or wind direction from known release sites. Because our approach is relatively simple and is resistant to systematic errors in the observational record, it has broad applicability beyond specific constituents and ecosystem types to identify a wide variety of limited-duration anomalies in flux tower data to ensure human health and industrial safety.

54 ENVIRONMENTAL SCIENCES↗

Adaptive learning-driven high-throughput synthesis of oxygen reduction reaction Fe–N–C electrocatalysts

Reducing human reliance on inefficient energy systems and fossil fuels has become more urgent due to the consequences of global climate change. However, traditional trial-and-error approaches have hampered our ability to accelerate the discovery and implementation of functional materials for efficient energy conversion devices, such as polymer electrolyte fuel cells (PEFCs). To address this, we develop an adaptive learning framework that integrates machine learning and state-of-the-art capabilities in high-throughput synthesis to achieve expedited optimization of iron-nitrogen-carbon PEFC oxygen reduction reaction (ORR) electrocatalysts. We use statistical inference, uncertainty quantification, and global optimization to build a computational design-of-experiment tool that identifies the optimum compositions to be investigated next to reduce the demands placed on experimental materials discovery. We benchmark the ability of the proposed strategy to discover optimum catalyst synthesis conditions in a six-dimensional search space when starting with a thirty-six-sample database. By following the adaptive learning strategy, we synthesize fourteen new catalysts from approximately ten billion unique compositions and discover four catalysts that outperform all original samples. The best machine learning-optimized catalyst is 33% more active than the highest-performing one in the initial database, showing an ORR activity seven times larger than those typically reported for the same class of materials.

36 MATERIALS SCIENCE↗

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An initial investigation of accuracy required for the identification of small molecules in complex samples using quantum chemical calculated NMR chemical shifts

The majority of primary and secondary metabolites in nature have yet to be identified, representing a major challenge for metabolomics studies that currently require reference libraries from analyses of authentic compounds. Using currently available analytical methods, complete chemical characterization of metabolomes is infeasible for both technical and economic reasons. For example, unambiguous identification of metabolites is limited by the availability of authentic chemical standards, which, for the majority of molecules, do not exist. Computationally predicted or calculated data are a viable solution to expand the currently limited metabolite reference libraries, if such methods are shown to be sufficiently accurate. For example, determining nuclear magnetic resonance (NMR) spectroscopy spectra in silico has shown promise in the identification and delineation of metabolite structures. Many researchers have been taking advantage of density functional theory (DFT), a computationally inexpensive yet reputable method for the prediction of carbon and proton NMR spectra of metabolites. However, such methods are expected to have some error in predicted 13 >C and 1 H NMR spectra with respect to experimentally measured values. This leads us to the question–what accuracy is required in predicted 13 C and 1 H NMR chemical shifts for confident metabolite identification? Using the set of 11,716 small molecules found in the Human Metabolome Database (HMDB), we simulated both experimental and theoretical NMR chemical shift databases. We investigated the level of accuracy required for identification of metabolites in simulated pure and impure samples by matching predicted chemical shifts to experimental data. We found 90% or more of molecules in simulated pure samples can be successfully identified when errors of 1 H and 13 C chemical shifts in water are below 0.6 and 7.1 ppm, respectively, and below 0.5 and 4.6 ppm in chloroform solvation, respectively. In simulated complex mixtures, as the complexity of the mixture increased, greater accuracy of the calculated chemical shifts was required, as expected. However, if the number of molecules in the mixture is known, e.g., when NMR is combined with MS and sample complexity is low, the likelihood of confident molecular identification increased by 90%.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Characterization and structure of the human lysine-2-oxoglutarate reductase domain, a novel therapeutic target for treatment of glutaric aciduria type 1

In humans, a single enzyme 2-aminoadipic semialdehyde synthase (AASS) catalyses the initial two critical reactions in the lysine degradation pathway. This enzyme evolved to be a bifunctional enzyme with both lysine-2-oxoglutarate reductase (LOR) and saccharopine dehydrogenase domains (SDH). Moreover, AASS is a unique drug target for inborn errors of metabolism such as glutaric aciduria type 1 that arise from deficiencies downstream in the lysine degradation pathway. While work has been done to elucidate the SDH domain structurally and to develop inhibitors, neither has been done for the LOR domain. Here, we purify and characterize LOR and show that it is activated by alkylation of cysteine 414 by N-ethylmaleimide. We also provide evidence that AASS is rate-limiting upon high lysine exposure of mice. Finally, we present the crystal structure of the human LOR domain. Our combined work should enable future efforts to identify inhibitors of this novel drug target.

59 BASIC BIOLOGICAL SCIENCES↗

An automated cluster surface scanning method for exploring reaction paths on metal-cluster surfaces

Metal-cluster surfaces present a wide variety of unique coordination environments. This complexity makes it difficult to manually probe the surface reactivity of such clusters. Here, we present a simple and automated method to systematically discover reaction pathways on cluster surfaces, based on the automated cluster surface scanning (ACSS) technique for mapping out potential energy surfaces. We showcase our method on 55-atom icosahedral Cu and Ag clusters, where we determine the activation energies of four elementary steps common in heterogeneous catalysis – hydrogen recombination (H* + H* → H 2 * + *), oxygen recombination (O* + O* → O 2 * + *), water formation (OH* + H* → H 2 O* + *), and CO oxidation (CO* + O* → CO 2 * + *) – with density functional theory calculations (DFT-PBE + D3). We show that the ACSS method requires significantly less human effort than the established manually performed climbing-image nudged elastic band (MP + CI-NEB) technique and locates transition states with comparable accuracy (root-mean-squared error of 0.10 eV) and similar computational cost. Rigorous sampling of the potential energy surface with the ACSS method allows one to locate all lowest-energy reaction pathways obtained via the MP+CI-NEB approach, as well as alternative pathways that one may have missed with the MP+CI-NEB approach due to the many possible pathways available on these clusters. The accuracy and efficiency afforded by the ACSS method could enable high-throughput exploration of the diverse reactivity of metal clusters.

36 MATERIALS SCIENCE↗

Term Matrix: a novel Gene Ontology annotation quality control system based on ontology term co-annotation patterns

Biological processes are accomplished by the coordinated action of gene products. Gene products often participate in multiple processes, and can therefore be annotated to multiple Gene Ontology (GO) terms. Nevertheless, processes that are functionally, temporally and/or spatially distant may have few gene products in common, and co-annotation to unrelated processes probably reflects errors in literature curation, ontology structure or automated annotation pipelines. We have developed an annotation quality control workflow that uses rules based on mutually exclusive processes to detect annotation errors, based on and validated by case studies including the three we present here: fission yeast protein-coding gene annotations over time; annotations for cohesin complex subunits in human and model species; and annotations using a selected set of GO biological process terms in human and five model species. For each case study, we reviewed available GO annotations, identified pairs of biological processes which are unlikely to be correctly co-annotated to the same gene products (e.g. amino acid metabolism and cytokinesis), and traced erroneous annotations to their sources. To date we have generated 107 quality control rules, and corrected 289 manual annotations in eukaryotes and over 52 700 automatically propagated annotations across all taxa.

59 BASIC BIOLOGICAL SCIENCES↗

Residual estimation for grid modification in wall-modeled large eddy simulation using unstructured high-order methods

Here, the accuracy and computational cost of a large eddy simulation are highly dependent on the computational grid. Building optimal grids manually from a priori knowledge is not feasible in most practical use cases; instead, solution-adaptive strategies can provide a robust and cost-efficient method to generate a grid with the desired accuracy. We adapt the residual estimation algorithm developed by Toosi and Larsson for Discontinuous Galerkin Spectral Elements Methods (DGSEM) to guide the grid-adaptation process. The core of the method is the computation of the estimated modeling residual using the polynomial basis functions used in DGSEM, and the averaging of the estimated residual over each element. The final method is assessed in multiple channel flow test cases and for the transonic flow over an airfoil, in both cases making use of mortar interfaces between elements with hanging nodes. The method is found to be robust and reliable, and to provide solutions on grids with significantly fewer elements at comparable accuracy compared to when using human-generated grids.

97 MATHEMATICS AND COMPUTING↗

Lab-Scale Cable-Driven Parallel Robot Prototype for Automated Prefabricated Component Manipulation

This paper presents the design and evaluation of a lab-scale cable-driven parallel robot (CDPR) developed as a flexible platform for automated installation of prefabricated components onto exterior building envelopes. Traditional manual installation methods for prefabricated components, which depend on scaffolding, cranes, cherry pickers, and verbal coordination, are not only labor-intensive and error-prone but also face significant limitations in dense urban environments due to site access constraints. To address these challenges, we developed a lab-scale CDPR platform capable of autonomously transporting building envelope components from a designated pickup zone to their target installation location, minimizing the need for human intervention. This study describes the system’s mechanical design, actuation architecture, real-time feedback system, and control strategy of the CDPR, and evaluates its performance in a laboratory environment. The robot’s actuation system uses torque control for end-effector manipulation. The robot’s real-time pose feedback comes from a construction-grade total station and a wireless inertial measurement unit (IMU), which together support precise end-effector control. Experimental results demonstrate the successful integration of the hardware, sensing, state estimation, and control subsystems. Preliminary tests showed that our lab-scale prototype can position the end effector with an error of less than 3 mm, which is a level of precision not previously achieved by existing CDPRs in construction applications. The key findings are twofold: (1) torque-only control is necessary but not sufficient for minimizing final pose error, and (2) incorporating real-time pose feedback can achieve the desired placement accuracy.

Liu, Yifang [Oak Ridge National Laboratory (ORNL),↗

Predicting bulge to total luminosity ratio of galaxies using deep learning

ABSTRACT We present a deep learning model to predict the r-band bulge-to-total luminosity ratio (B/T) of nearby galaxies using their multiband JPEG images alone. Our Convolutional Neural Network (CNN) based regression model is trained on a large sample of galaxies with reliable decomposition into the bulge and disc components. The existing approaches to estimate the B/T ratio use galaxy light-profile modelling to find the best fit. This method is computationally expensive, prohibitively so for large samples of galaxies, and requires a significant amount of human intervention. Machine learning models have the potential to overcome these shortcomings. In our CNN model, for a test set of 20 000 galaxies, 85.7 per cent of the predicted B/T values have absolute error (AE) less than 0.1. We see further improvement to 87.5 per cent if, while testing, we only consider brighter galaxies (with r-band apparent magnitude <17) with no bright neighbours. Our model estimates the B/T ratio for the 20 000 test galaxies in less than a minute. This is a significant improvement in inference time from the conventional fitting pipelines, which manage around 2–3 estimates per minute. Thus, the proposed machine learning approach could potentially save a tremendous amount of time, effort, and computational resources while predicting B/T reliably, particularly in the era of next-generation sky surveys such as the Legacy Survey of Space and Time (LSST) and the Euclid sky survey which will produce extremely large samples of galaxies.

Grover, Harsh (ORCID:0000000321338142)↗

A multi-scale cognitive interaction model of instrument operations at the Linac Coherent Light Source

The Linac Coherent Light Source (LCLS) is the world’s first x-ray free electron laser. It is a scientific user facility operated by the SLAC National Accelerator Laboratory, at Stanford, for the U.S. Department of Energy. As beam time at LCLS is extremely valuable and limited, experimental efficiency—getting the most high quality data in the least time—is critical. Our overall project employs cognitive engineering methodologies with the goal of improving experimental efficiency and increasing scientific productivity at LCLS by refining experimental interfaces and workflows, simplifying tasks, reducing errors, and improving operator safety and stress. Here, in this study, we describe a multi-agent, multi-scale computational cognitive interaction model of instrument operations at LCLS. Our model simulates the aspects of human cognition at multiple cognitive and temporal scales, ranging from seconds to hours, and among agents playing multiple roles, including instrument operator, real time data analyst, and experiment manager. The model can roughly predict impacts stemming from proposed changes to operational interfaces and workflows. Example results demonstrate the model’s potential in guiding modifications to improve operational efficiency. We discuss the implications of our effort for cognitive engineering in complex experimental settings and outline future directions for research. The model is open source, and the videos of the supplementary material provide extensive detail.

47 OTHER INSTRUMENTATION↗

From Text to Maps: LLM-Driven Extraction and Geotagging of Epidemiological Data

Epidemiological datasets are essential for public health analysis and decision-making, yet they remain scarce and often difficult to compile due to inconsistent data formats, language barriers, and evolving political boundaries. Traditional methods of creating such datasets involve extensive manual effort and are prone to errors in accurate location extraction. To address these challenges, we propose utilizing large language models (LLMs) to automate the extraction and geotagging of epidemiological data from textual documents. Our approach significantly reduces the manual effort required, limiting human intervention to validating a subset of records against text snippets and verifying the geotagging reasoning, as opposed to reviewing multiple entire documents manually to extract, clean, and geotag. Additionally, the LLMs identify information often overlooked by human annotators, further enhancing the dataset’s completeness. Our findings demonstrate that LLMs can be effectively used to semi-automate the extraction and geotagging of epidemiological data, offering several key advantages: (1) comprehensive information extraction with minimal risk of missing critical details; (2) minimal human intervention; (3) higher-resolution data with more precise geotagging; and (4) significantly reduced resource demands compared to traditional methods.

Harrod, Karly↗

Comparison of Error Rate Depending on Operator Expertise and Simulator Complexity

This paper analyzes operator's error rate from experiments, depending on the expertise and simulator complexity. This study uses the Rancor Microworld, a simplified simulator developed by INL, and the Compact Nuclear Simulator (CNS), a less simplified simulator developed by the Korea Atomic Energy Research Institute (KAERI). The error rates were measured from simulation data of a total of 72 participants, and the collected error rate data were analyzed using analysis of variance (ANOVA) test.

99 GENERAL AND MISCELLANEOUS↗

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic compendium of 189,680 DNA viruses from the human gut microbiome

Bacteriophages have important roles in the ecology of the human gut microbiome but are under-represented in reference databases. To address this problem, we assembled the Metagenomic Gut Virus catalogue that comprises 189,680 viral genomes from 11,810 publicly available human stool metagenomes. Over 75% of genomes represent double-stranded DNA phages that infect members of the Bacteroidia and Clostridia classes. Based on sequence clustering we identified 54,118 candidate viral species, 92% of which were not found in existing databases. The Metagenomic Gut Virus catalogue improves detection of viruses in stool metagenomes and accounts for nearly 40% of CRISPR spacers found in human gut Bacteria and Archaea. We also produced a catalogue of 459,375 viral protein clusters to explore the functional potential of the gut virome. This revealed tens of thousands of diversity-generating retroelements, which use error-prone reverse transcription to mutate target genes and may be involved in the molecular arms race between phages and their bacterial hosts.

59 BASIC BIOLOGICAL SCIENCES↗

An Approach to Realize Generalized Optimal Motion Primitives Using Physics Informed Neural Networks

Autonomous manipulation is a challenging problem in field robotics due to uncertainty in object properties, constraints, and coupling phenomenon with robot control systems. Humans learn motion primitives over time to effectively interact with the environment. We postulate that autonomous manipulation can be enabled by basic sets of motion primitives as well, but do not necessitate mimicking human motion primitives. Here, this work presents an approach to generalized optimal motion primitives using physics-informed neural networks. Our simulated and experimental results demonstrate that optimality is notionally maintained where the mean maximum observed final position percent error was 0.564% and the average mean error for all the trajectories was 1.53%. These results indicate that notional generalization is attained using a physics-informed neural network approach that enables near optimal real-time adaptation of primitive motion profiles.

97 MATHEMATICS AND COMPUTING↗

Harmonizing direct and indirect anthropogenic land carbon fluxes indicates a substantial missing sink in the global carbon budget since the early 20th century

Inconsistencies in the calculation of the two anthropogenic land flux terms of the global carbon cycle are investigated. The two terms—the direct anthropogenic flux (caused by direct human disturbance in anthromes, currently a carbon source to the atmosphere) and the indirect anthropogenic flux (caused indirectly by human activities that lead to global change and affecting all biomes, currently an atmospheric carbon sink)—are typically calculated independently, resulting in inconsistent underlying assumptions. We harmonize the estimation of the two anthropogenic land flux terms by incorporating previous estimates of these inconsistencies. We recalculate the global carbon budget (GCB) and apply change-point analysis to the cumulative budget imbalance. Cumulative over 1850–2018 (1959–2018), harmonization results in a 13% lesser (4% greater) land use source from anthromes and a 20% (23%) lesser land sink. This recalculation yields a greater non-closure of the GCB, indicating a missing carbon sink averaging 0.65 Pg C year -1 since the early 20th century. The imbalance likely results from a combination of method discontinuity and structural errors in the assessment of the direct anthropogenic land use flux, greater ocean carbon uptake, structural errors in land models, and in how these land terms are quantified for the budget. We caution against overconfidence in considering the GCB a solved problem and recommend further study of methodological discontinuities in budget terms. We strongly recommend studies that quantify the direct and indirect anthropogenic land fluxes simultaneously to ensure consistency, with a deeper understanding of human disturbance and legacy effects in anthromes.

54 ENVIRONMENTAL SCIENCES↗