Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “keyword”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Depletion of atmospheric neutrino fluxes from parton energy loss

The phenomenon of fully coherent energy loss (FCEL) in the collisions of protons on light ions affects the physics of cosmic ray air showers. As an illustration, we address two closely related observables: hadron production in forthcoming proton-oxygen collisions at the LHC, and the atmospheric neutrino fluxes induced by the semileptonic decays of hadrons produced in proton-air collisions. In both cases, a significant nuclear suppression due to FCEL is predicted. The conventional and prompt neutrino fluxes are suppressed by ~10...25% in their relevant neutrino energy ranges. Previous estimates of atmospheric neutrino fluxes should be scaled down accordingly to account for FCEL. Keywords: Atmospheric neutrinos, Parton energy loss

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format

The layout of portable document format (PDF) files is constant to any screen, and the metadata therein are latent, compared to mark-up languages such as HTML and XML. No semantic tags are usually provided, and a PDF file is not designed to be edited or its data interpreted by software. However, data held in PDF files need to be extracted in order to comply with opensource data requirements that are now government-regulated. In the chemical domain, related chemical and property data also need to be found, and their correlations need to be exploited to enable data science in areas such as data-driven materials discovery. Such relationships may be realized using text-mining software such as the “chemistry-aware” natural-language-processing tool, ChemDataExtractor; however, this tool has limited data-extraction capabilities from PDF files. This study presents the PDFDataExtractor tool, which can act as a plug-in to ChemDataExtractor. It outperforms other PDF-extraction tools for the chemical literature by coupling its functionalities to the chemical-named entityrecognition capabilities of ChemDataExtractor. The intrinsic PDF-reading abilities of ChemDataExtractor are much improved. The system features a template-based architecture. This enables semantic information to be extracted from the PDF files of scientific articles in order to reconstruct the logical structure of articles. While other existing PDF-extracting tools focus on quantity mining, this template-based system is more focused on quality mining on different layouts. PDFDataExtractor outputs information in JSON and plain text, including the metadata of a PDF file, such as paper title, authors, affiliation, email, abstract, keywords, journal, year, document object identifier (DOI), reference, and issue number. With a self-created evaluation article set, PDFDataExtractor achieved promising precision for all key assessed metadata areas of the document text.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Systems Engineering and Analysis in Support of a US Federal Staging Facility for UNF

The US Department of Energy Office of Nuclear Energy (DOE-NE) Office of Spent Fuel and High-Level Waste Disposition is examining a set of system options and conducting supporting analyses to inform the development of an integrated waste management system, which may include one or more federal staging facilities (FSFs) for used nuclear fuel (UNF ) sited using a collaborative siting process. This paper focuses on the ongoing activities in two systems engineering and analysis work areas: (1) data and tools development, validation, and maintenance and (2) systems engineering execution. Within the first work area, the STANDARDS 5.0 UNF data and analysis tool, formerly known as UNF-ST&DARDS, is being developed as a foundational resource to assist in the management of UNF data. It has the key capability to model UNF throughout the entire back end of the fuel cycle. STANDARDS also includes several compatible analysis tools for the time-dependent characterization of UNF and related systems by interfacing with the SCALE code system for nuclear analysis and COBRA-SFS for thermal analysis. Also, within the data and tools area is the Next Generation System Analysis Model (NGSAM), which is an agent-based simulation software tool expressly designed to be capable of modeling the waste management system, including the transportation of UNF to and from a FSF. NGSAM has been developed to enable informed decision-making by providing the capability to analyze various potential system options for the management of UNF and high-level radioactive waste. Finally, in the systems engineering execution area, the team has begun to apply a disciplined systems engineering approach at the system level along with supporting analysis to guide the development of the FSF project requirements (including associated transportation infrastructure). Systems engineering principles and practices and their adaptation/application to design and development activities will ensure that the waste management system is effectively implemented as work proceeds. Other activities include investigating the implications of changes in various assumptions and parameters related to waste management systems, such as UNF acceptance rates, receipt logic, facility capacities and capabilities, use of standardized canisters, and different assumed facility operation start dates. Keywords: federal staging facility (FSF), used nuclear fuel (UNF), integrated waste management (IWM) system, Next Generation System Analysis Model (NGSAM), STANDARDS, systems engineering

Joseph, Robert↗

Bibliometric review and recent advances in total scattering pair distribution function analysis: 21 years in retrospect

Global research activities have been driven by the quest to develop and characterize novel materials for technological advancements. The total scattering pair distribution function (TSPDF) is a powerful and versatile characterization technique for examining the structural details of diverse complex materials including liquid, amorphous, disordered crystalline, and nanostructured materials. Thus, it is critical to keep track of research progress, identify research gaps, and future research directions of the application of the TSPDF technique in materials development and discovery. In this work, a bibliometric analysis of literature regarding the TSPDF technique between 2000 and 2021 was conducted using datasets retrieved from the Web of Science database. The research trends based on publication outputs, research subject distribution, co-authorships among institutions, countries/regions, co-citation of referenced sources, and keyword co-occurrence are evaluated and discussed herein. The impact of the TSPDF technique is projected to increase due to its importance in probing emerging functional materials, and the advances in specialized facilities and instrumentation among the scientific communities engaged with it. Finally, current and emerging research hotspots related to TSPDF technique such as catalysis, computer modeling and simulation, pharmaceutics, machine learning, hydrogen storage, battery materials, and layered structured materials are also identified and discussed.

36 MATERIALS SCIENCE↗

Mapping the structural, magnetic and electronic behavior of (Eu 1-x Ca x ) 2 Ir 2 O 7 across a metal-insulator transition

In this study, we employ bulk electronic properties characterization and x-ray scattering/spectroscopy techniques to map the structural, magnetic and electronic properties of (Eu 1-x Ca x ) 2 Ir 2 O 7 as a function of Ca-doping. As expected, the metal-insulator transition temperature, T-MIT, decreases with Ca-doping until a metallic state is realized down to 2 K. In contrast, T-AFM becomes decoupled from the MIT and (likely short-range) AFM order persists into the metallic regime. This decoupling is understood as a result of the onset of an electronically phase separated state, the occurrence of which seemingly depends on both synthesis method and rare earth site magnetism. PDF analysis suggests that electronic phase separation occurs without accompanying chemical phase segregation or changes in the short-range crystallographic symmetry while synchrotron x-ray diffraction confirms that there is no change in the long-range crystallographic symmetry. X-ray absorption measurements confirm the $J_{eff}$ = 1/2 character of (Eu 1-x Ca x ) 2 Ir 2 O 7 . Surprisingly these measurements also indicate a net electron doping, rather than the expected hole doping, indicating a compensatory mechanism. Lastly, XMCD measurements show a weak Ir magnetic polarization that is largely unaffected by Ca-doping. Keywords: term, term, term.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

RNAcentral 2021: secondary structure integration, improved sequence search and new member databases

RNAcentral is a comprehensive database of non-coding RNA (ncRNA) sequences that provides a single access point to 44 RNA resources and >18 million ncRNA sequences from a wide range of organisms and RNA types. RNAcentral now also includes secondary (2D) structure information for >13 million sequences, making RNAcentral the world’s largest RNA 2D structure database. The 2D diagrams are displayed using R2DT, a new 2D structure visualization method that uses consistent, reproducible and recognizable layouts for related RNAs. The sequence similarity search has been updated with a faster interface featuring facets for filtering search results by RNA type, organism, source database or any keyword. This sequence search tool is available as a reusable web component, and has been integrated into several RNAcentral member databases, including Rfam, miRBase and snoDB. To allow for a more fine-grained assignment of RNA types and subtypes, all RNAcentral sequences have been annotated with Sequence Ontology terms. The RNAcentral database continues to grow and provide a central data resource for the RNA community. RNAcentral is freely available at https://rnacentral.org.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing Sub-Cohorts via Data Normalization and Representation Learning

The process of identifying a cohort of interest is a very challenging task. It requires manually inspecting many patient records of complex structure that might include medical coding errors and missing data. This paper presents a computational pipeline for refining the process of cohort selection based on medical concepts recorded in the electronic health records (EHRs). The pipeline extracts EHR data for a given cohort and normalizes this data using standard vocabularies. Then a stacked denoising autoencoder is used to embed the normalized patient vectors in a low dimensional space, where the patients are subsequently clustered into sub-cohorts. The goal is to represent the cohort in a standard format and abstract variants of sub-populations. As a use-case, we applied the pipeline to 1.8 million Veterans diagnosed with major depressive disorder (MDD), and identified four meaningful sub-cohorts using the features learned by the autoencoder. Then, each sub-cohort was explored using a set of keywords for interpretation.

Rush III, Everett↗

ORGANIC RANKINE CYCLE TURBINE AND HEAT EXCHANGER SIZING FOR LIQUID AIR COMBINED CYCLE

Cryogenic energy storage offers several opportunities to design turbomachinery and other equipment for novel cycles. This paper presents the design and analysis of turbomachinery and heat exchangers for an Organic Rankine Cycle (ORC) subsystem for a hybrid energy storage concept. The Liquid Air Combined Cycle is an energy storage system that stores air at cryogenic conditions at times with high variable renewable energy to be dispatched along with a gas turbine to recover the exhaust heat. In order to re-vaporize the air, the liquid air is coupled with an ORC as an additional bottoming cycle. The ORC turbine is expected to expand the fluid with a pressure ratio of nearly 30 and a flow rate of approximately 45 kg/s. Sizing calculations for both a radial and axial turbine solution were performed over a range of speeds and stages to determine the optimal design point. The results show that either an axial (8- or 9-stage) or radial (four stages at two shaft speeds) turbine are capable of handling the pressure ratios. Further trades of the two configurations would be required to determine the best option. The ORC system also incorporates five heat exchangers to distribute heat, vaporize the liquid air, or recover exhaust heat from the gas turbine. Three heat exchangers were analyzed to understand the size of heat exchangers and pressure drop for the overall system. Different types of heat exchangers were explored for the different purposes, including plate-fin heat exchangers, gasketed plate heat exchangers and shell-in-tube heat exchangers. It was determined that the ORC recuperator, liquid-air vaporizer, and vaporized air pre-heater would be counter-flow heat exchangers using a gasketed plate design. Keywords: Energy Storage, Liquid Air Energy Storage, Organic Rankine Cycle

Pryor, Owen↗

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

Web of Registries Search (WoRS) v1.0.0

The Web of Registries Search (WoRs) is a web based software application that enables users to search for publicly available biological parts using keywords or sequence fragments. For the initial version (1.0.0) of the application, WoRS targets 10 sources of biological part data: the GenBank NIH genetic sequence database (https://www.ncbi.nlm.nih.gov/genbank/), the iGem parts registry (parts.igem.org), the Addgene plasmid repository (https://www.addgene.org/), and 7 Inventory of Composable Elements (ACS Synbio, JGI, JBEI, JBEI Public, ABF, SynBerc, ABF Public) registry instances. WoRs has built-in automated web scrapers which extract data from sources that do not have a public or well-defined application programming interface (API). They extract as much public data as they can find and create a searchable index to speed up searches. Included in the indexed information is the source of the information.

Plahar, Hector↗

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector↗

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS↗

BIN1 protein isoforms are differentially expressed in astrocytes, neurons, and microglia: neuronal and astrocyte BIN1 are implicated in tau pathology

Background: Identified as an Alzheimer’s disease (AD) susceptibility gene by genome wide-association studies, BIN1 has 10 isoforms that are expressed in the Central Nervous System (CNS). The distribution of these isoforms in different cell types, as well as their role in AD pathology still remains unclear. Methods: Utilizing antibodies targeting specific BIN1 epitopes in human post-mortem tissue and analyzing mRNA expression data from purified microglia, we identified three isoforms expressed in neurons and astrocytes (isoforms 1, 2 and 3) and four isoforms expressed in microglia (isoforms 6, 9, 10 and 12). The abundance of selected peptides, which correspond to groups of BIN1 protein isoforms, was measured in dorsolateral prefrontal cortex, and their relation to neuropathological features of AD was assessed. Results: Peptides contained in exon 7 of BIN1’s N-BAR domain were found to be significantly associated with ADrelated traits and, particularly, tau tangles. Decreased expression of BIN1 isoforms containing exon 7 is associated with greater accumulation of tangles and subsequent cognitive decline, with astrocytic rather than neuronal BIN1 being the more likely culprit. These effects are independent of the BIN1 AD risk variant. Conclusions: Exploring the molecular mechanisms of specific BIN1 isoforms expressed by astrocytes may open new avenues for modulating the accumulation of Tau pathology in AD. Keywords: BIN1 isoforms, Alzheimer’s disease, Amyloid, Tau, Microglia, Astrocytes, Neurons

Taga, Mariko↗

Earth microbial co-occurrence network reveals interconnection pattern across microbiomes

Microbial interactions shape the structure and function of microbial communities; microbial association networks in specic environments have been widely developed to explore these complex systems, but their wired pattern across microbiomes in various environments at the global scale remains unexplored. Here we have inferred an Earth microbial association network from a communal catalogue with 23,595 samples and 12,646 exact sequence variants from 14 environments in the Earth Microbiome Project dataset. Results: This non-random scale-free Earth microbial association network consisted of 8 taxonomy distinct modules linked with dierent environments, which featured environment specic microbial associations. Dierent topological features of subnetworks inferred from datasets trimmed into uniform size indicate distinct association patterns in the microbiomes of various environments. The proportions of specialist edges, which ranged from 43.0% to 65.7%, highlight that environmental specic microbial associations are essential features of microbiomes in various environments. Based on edge-overlap similarity, the microbiomes of various environments were clustered into two groups, which were mainly bridged by the microbiomes of plant and animal surface. Acidobacteria Gp2 and Nisaea were identied as hubs in most of subnetworks. Negative edges proportions ranged from 1.9% in the soil subnetwork to 48.9% the non-saline surface subnetwork, suggesting various environments experience distinct intensities of competition or niche dierentiation. Conclusion: This investigation provides a new resource for examining Earth microbial association patterns across environments and emphasizes the network perspective for comprehensively understanding unique microbiome features. Keywords: Association pattern; Earth microbiomes; Genelist edges; Network hubs; Negative associations; Specialist edges; Topological properties

59 BASIC BIOLOGICAL SCIENCES↗

Calibration of V-Notch and Compound Weirs for Subsurface Drainage Water Level Control Structures

Highlights Accurate discharge estimation is important when evaluating edge-of-field conservation practices. V-notch weir equations were developed for three sizes of subsurface drainage water level control structures. Compound weir equations were developed for subsurface drainage water level control structures. The compound weir equation accurately estimates discharge for flows within and overtopping the V-notch. Abstract.Numerous edge-of-field conservation practices use subsurface drainage water level control structures to monitor water levels and estimate discharge. In a control structure, procedures for calculating discharge when flow depth (head) exceeds the V-notch depth and overflows in the rectangular portion of the compound weir (CW) are ambiguous. In this study, we developed calibration equations for V-notch weirs in Agri Drain inline water level control structures of different sizes for flows within the V-notch and overtopping flow events. The discharge equation for overtopping events (Q CW , L s -1 ) was determined as: Q CW = a 1 (h b1 -h 1 b1 )+a 2 (W e -W v )h 1 b2 , where h and h 1 are heads above vertex/bottom and top of V-notch (cm), respectively, W is the effective crest width of rectangular weir (cm), W v is the top width of V-notch (cm), a 1 and b 1 are parameters for V-notch weir obtained by calibration, and a 2 and b 2 are calibration parameters for rectangular weir obtained from literature. Results were compared with a weir equation available in the literature (Q V+R ), which combines a V-notch equation with a head equal to V-depth and a rectangular weir equation for flow above V-depth. Discharge at overflow was estimated with high accuracy with Q CW, whereas Q V+R underestimated discharge (e.g., PBIAS of 0.67% vs. 17.82% for a 15.2 cm structure). An example using Q V+R resulted in a 14% lower annual estimation of nitrate-N load diverted to a saturated buffer than Q CW due to underestimation of drainage discharge during overflow events. Results suggest that the developed equation (Q CW ) accurately estimates discharge and will thus improve the estimated N load compared to Q V+R . Keywords: Compound weir, Flow monitoring, Subsurface drainage, V-notch weir, Water level control structure, Weir calibration.

Agriculture↗

Effectiveness of Residue and Tillage Management on Runoff Pollutant Reduction from Agricultural Areas

Highlights No-till and no-till residue systems were effective in reducing runoff particulate and total nutrients but increased dissolved nutrients. Maintaining >30% residue cover reduced most runoff constituents, irrespective of no-till or tillage. No-till-residue prevented runoff nutrient losses and benefitted farm revenue by avoiding tillage. Abstract. Reduced tillage management conservation practices (No-till and Reduced-till) are widely adopted in agriculture; however, understanding their overall effectiveness for water quality protection is challenging. A meta-analysis was conducted to understand and quantify the effectiveness of residue and tillage management on runoff, sediment, and nutrient losses from agricultural fields. Annual runoff and the associated sediment, and nutrient (nitrogen and phosphorus) loads were compiled from 60 peer reviewed research articles published across the United States and Canada. A total of 1575 site-years of data were categorized into tillage (<30% surface cover), no-tillage (<30% surface cover), tillage with residue (>30% surface cover), no-tillage with residue (>30% surface cover), and pasture management. No-tillage, no-tillage-residue, and tillage-residue managements were evaluated for their effectiveness in reducing runoff, nutrients, and sediment loads compared to tillage. Synthesized and surveyed corn yield data were used to evaluate the economic cost effectiveness of no-tillage-residue management with respect to tillage. Across the site years (1968-2019) studied, median runoff depth for no-tillage and no-tillage-residue were 84% and 70% greater than tillage and tillage-residue management, respectively. No-tillage-residue management had up to 86% less sediment losses than tillage systems, on average, for both >30% and <30% surface cover. No-tillage-residue management was most effective, with a positive performance effectiveness of 65% to 90% in controlling sediments, particulate, and total nutrient losses in runoff compared to tillage. Cost effectiveness analysis revealed the benefits of no-tillage-residue management in reducing nutrient loads and increasing net-farm revenue by avoiding tillage operational costs. Except for dissolved phosphorus, no-tillage-residue management cost effectiveness for sediments and nutrient loads ranged from negative $6 to negative $102 per every Mg or kg of load reduction, indicating it had both economic and environmental benefits compared to tillage management. Overall, these results indicate that over the long-term, no-tillage and tillage, combined with greater than 30% residue cover, can effectively reduce sediment and nutrient losses. This work highlights the importance of crop residues on the soil surface to reduce runoff losses, even in no-tillage systems. Keywords: Conservation tillage, No-tillage, Residue cover, Tillage, Water quality.

Agriculture↗