Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Keywords”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Semantic Search with Sentence-BERT for Design Information Retrieval

Managing and referencing design knowledge is a critical activity in the design process. However, reliably retrieving useful knowledge can be a frustrating experience for users of knowledge management systems due to inherent limitations of standard keyword-based searches. In this research, we consider the task of retrieving relevant lessons learned from the NASA Lessons Learned Information System (LLIS). To this end, we apply a state-of-the-art natural language processing (NLP) technique for information retrieval (IR): semantic search with sentence-BERT, which is a modification of a Bidirectional Encoder Representations from Transformers (BERT) model that uses siamese and triplet network architectures to obtain semantically meaningful sentence embeddings. While the pre-trained sBERT model performs well out-of-the-box, we further fine-tune the model on data from the LLIS so that it learns on design engineering-relevant vocabulary. We quantify the improvement in query results using both standard sBERT and fine-tuned sBERT over a keyword search. Our use case throughout the paper is to use queries related to specific requirements from a NASA project. Fine tuning the sBERT model on LLIS data yields a mean average precision (MAP) of 0.807 on queries based on information needs from a real NASA project. Results indicate that applying state-of-the-art natural language processing techniques, especially when finetuned using engineering data, to design information retrieval tasks shows significant promise in modernizing design knowledge management systems.

Hannah S. Walsh↗

Semantic Search with Sentence-BERT for Design Information Retrieval

Managing and referencing design knowledge is a critical activity in the design process. However, reliably retrieving useful knowledge can be a frustrating experience for users of knowledge management systems due to inherent limitations of standard keyword-based searches. In this research, we consider the task of retrieving relevant lessons learned from the NASA Lessons Learned Information System (LLIS). To this end, we apply a state-of-the-art natural language processing (NLP) technique for information retrieval: semantic search with sentence-BERT, which is a modification of a Bidirectional Encoder Representations from Transformers (BERT) model that uses siamese and triplet network architectures to obtain semantically meaningful sentence embeddings. While the pretrained sBERT model shows excellent out-of-the-box performance, we further fine-tune the model on data from the LLIS so that it learns on design engineering-relevant vocabulary. We quantify the improvement in query results using both standard sBERT and fine-tuned sBERT over the LLIS’s built-in keyword search. Additionally, we demonstrate a use case for the query system by searching for lessons learned relevant to specific requirements from a NASA project as part of a broader knowledge management and retrieval system. Results indicate that applying state-of-the-art natural language processing techniques, especially when fine-tuned using engineering data, to design information retrieval tasks shows significant promise in modernizing design knowledge management systems.

Hannah S. Walsh↗

Informing NLP Learning Tasks by Tracking User Features: An ASRS Use Case using Kaona

There has been growing interest in utilizing natural language processing (NLP) algorithms in Aviation Safety. This interest has extended to leveraging the decades of records publicly available on the Aviation Safety Reporting System (ASRS). While related literature has given more emphasis in lessons learned from the narratives, our prior work has focused on using NLP to support narrative search in the ASRS. Specifically, we evaluated if the use of alternative search mechanisms to keyword search, such as the retrieval of related narratives even without matching keywords could improve narrative discovery. A difficulty in experimenting alternative search mechanisms in any information retrieval task is the lack of ground truth. To address this limitation, we propose Kaona, a lightweight interface which enables the prototyping of alternative search retrieval tasks, by tracking user experience both explicitly (user-specified feedback), or implicitly (user navigation through interface affordances). Differently from distracting requests for feedback during user navigation, Kaona collects explicit feedback from users by mapping them to affordances which support the user workflow, while obtaining ground truth information for learning tasks.

human-computer-interaction↗

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

SURFACE FINISHING AND ELECTROLESS NICKEL PLATING OF ADDITIVELY MANUFACTURED (AM) METAL COMPONENTS

This study investigates the application of electroless nickel deposition on additively manufactured stainless steel samples. Current additive manufacturing (AM) technologies produce metal components with a rough surface. Rough surfaces generally exhibit fatigue characteristics, increasing the probability of initiating a crack or fracture to the printed part. For this reason, the direct use of as-produced parts in a finished product cannot be actualized, which presents a challenge. Post-processing of the AM parts is therefore required to smoothen the surface. This study analyzes chempolish (CP) and electropolish (EP) surface finishing techniques for post-processing AM stainless steel components CP has a great advantage in creating uniform, smooth surfaces regardless of size or part geometry EP creates an extremely smooth surface, which reduces the surface roughness to the sub-micrometer level. In this study, we also investigate nickel deposition on EP, CP, and as-built AM components using electroless nickel solutions. Electroless nickel plating is a method of alloy treatment designed to increase manufactured component's hardness and surface resistance to the unrelenting environment. The electroless nickel plating process is more straightforward than its counterpart electroplating.. We use low-phosphorus (2-5% P), medium-phosphorus (6-9% P), and high-phosphorus (10-13% P). These Ni deposition experiments were optimized using the L9 Taguchi design of experiments (TDOE), which compromises the prosperous content in the solution, surface finish, plane of the geometry, and bath temperature. The pre-and post-processed surface of the AM parts was characterized by KEYENCE Digital MicroscopeVHX-7000 and Phenom XL Desktop SEM. The experimental results show that electroless nickel deposition produces uniform Ni coating on the additively manufactured components up to 20 μm per hour. Mechanical properties of as-built and Ni coated AM samples were analyzed by applying a standard 10 N scratch test. Nickel coated AM samples were up to two times scratch resistant compared to the as-built samples. This study suggests electroless nickel plating is a robust viable option for surface hardening and finishing AM components for various applications and operating conditions.

Keywords: additive manufacturing, fatigue, chempol↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

PDB ‐101: Educational resources supporting molecular explorations through biology and medicine

Abstract The Protein Data Bank (PDB) archive is a rich source of information in the form of atomic‐level three‐dimensional (3D) structures of biomolecules experimentally determined using macromolecular crystallography, nuclear magnetic resonance (NMR) spectroscopy, and electron microscopy (3DEM). Originally established in 1971 as a resource for protein crystallographers to freely exchange data, today PDB data drive research and education across scientific disciplines. In 2011, the online portal PDB‐101 was launched to support teachers, students, and the general public in PDB archive exploration ( pdb101.rcsb.org ). Maintained by the Research Collaboratory for Structural Bioinformatics PDB, PDB‐101 aims to help train the next generation of PDB users and to promote the overall importance of structural biology and protein science to nonexperts. Regularly published features include the highly popular Molecule of the Month series, 3D model activities, molecular animation videos, and educational curricula. Materials are organized into various categories (Health and Disease, Molecules of Life, Biotech and Nanotech, and Structures and Structure Determination) and searchable by keyword. A biennial health focus frames new resource creation and provides topics for annual video challenges for high school students. Web analytics document that PDB‐101 materials relating to fundamental topics (e.g., hemoglobin, catalase) are highly accessed year‐on‐year. In addition, PDB‐101 materials created in response to topical health matters (e.g., Zika, measles, coronavirus) are well received. PDB‐101 shows how learning about the diverse shapes and functions of PDB structures promotes understanding of all aspects of biology, from the central dogma of biology to health and disease to biological energy.

Zardecki, Christine↗

Silica–Derived Nanostructured Electrode Materials for ORR, OER, HER, CO 2 RR Electrocatalysis, and Energy Storage Applications: A Review**

Silica-derived nanostructured catalysts (SDNCs) are a class of materials synthesized using nanocasting and templating techniques, which involve the sacrificial removal of a silica template to generate highly porous nanostructured materials. The surface of these nanostructures is functionalized with a variety of electrocatalytically active metal and non-metal atoms. SDNCs have attracted considerable attention due to their unique physicochemical properties, tunable electronic configuration, and microstructure. These properties make them highly efficient catalysts and promising electrode materials for next generation electrocatalysis, energy conversion, and energy storage technologies. The continued development of SDNCs is likely to lead to new and improved electrocatalysts and electrode materials. This review article provides a comprehensive overview of the recent advances in the development of SDNCs for electrocatalysis and energy storage applications. It analyzes 337,061 research articles published in the Web of Science (WoS) database up to December 2022 using the keywords “silica”, “electrocatalysts”, “ORR”, “OER”, “HER”, “HOR”, “CO 2 RR”, “batteries”, and “supercapacitors”. The review discusses the application of SDNCs for oxygen reduction reaction (ORR), oxygen evolution reaction (OER), hydrogen evolution reaction (HER), carbon dioxide reduction reaction (CO 2 RR), supercapacitors, lithium-ion batteries, and thermal energy storage applications. It concludes by discussing the advantages and limitations of SDNCs for energy applications.

25 ENERGY STORAGE↗

Workforce planning: a review of methodologies

Workforce planning deals with determining the number of employees and associated skills necessary to meet the future operational needs of an organization. A workforce system consists of six elements: recruitment, attrition, promotion, training, retention, and scheduling. Historically, several workforce modeling and analysis methodologies have been developed to capture these elements. This paper reviews the results of workforce and manpower models published within peer-reviewed literature between 1959 and 2021 to provide an in-depth analysis of current models. The focus of this review is on analytical, simulation, and empirical models found in literature that were collected based on a citation requirement and keyword search criteria. Results demonstrate the trends in workforce modeling research and discuss the common uses of each model type and the advantages/disadvantages related to each model. Based on the common attributes of workforce systems, the discussion focuses on the most frequently used model type for each element and the best use for each model. Lastly, recommendations are made for the development of workforce models that allow the most comprehensive view of the workforce systems of the future.

42 ENGINEERING↗

Visual Brick model authoring tool for building metadata standardization

In this study, the Brick ontology is a unified semantic metadata standard for building assets and their relationships, serving as a key enabler for effective interoperability and automation of building systems and analytics. However, creating a Brick model, in other words, standard semantic metadata based on the Brick ontology for a building dataset, can be a complex task. This paper presents two case studies of the creation of Brick models for real-world residential and commercial building datasets, highlighting the challenges during the Brick model creation process. Additionally, the paper introduces VizBrick, an interactive authoring tool for creating semantic building metadata. VizBrick facilitates the creation of Brick models by providing an intuitive visual interface and interactive capabilities, such as keyword search, automatic mapping suggestions, and recommendations. The use of VizBrick is shown to significantly reduce the time and effort required during the Brick model creation process.

42 ENGINEERING↗

Estimated capital costs of fish exclusion technologies for hydropower facilities

Hydropower is a reliable source of renewable energy, and its future expansion is likely to be in the form of either smaller new stream development (NSD) projects or powering existing non-powered dams. Thresholds for entrainment risk to fish and the requirements for fish exclusion at hydropower facilities often differ depending on the species involved, the characteristics of the facility, and the goals of stakeholders, but little quantitative information is present within the literature regarding the specific costs of fish exclusion measures. Cost data associated with protection, mitigation, and enhancement (PM&E) measures related to positive barrier screening were identified using keyword searches of an existing environmental mitigation cost data set and manual extraction from regulatory licensing documents available in the Federal Energy Regulatory Commission (FERC) eLibrary. This approach yielded a total of 50 p.m.&E mitigation measures with estimated capital construction costs pertaining to positive barrier screens and represented <10% of the 171 total FERC project dockets available in the data set. These data were highly skewed toward conventional relicensing projects, as <7% were associated with NSD projects. Results indicate highly variable costs are associated with fish screening, with flow-normalized costs one to two orders of magnitude higher for screening with the highest exclusion capability (≤0.09 in. spacing) compared with coarser screening (1–2 in.). These data provide an initial baseline for estimating exclusion costs for hydropower development and may help developers consider options for more fish-friendly generation technologies, though gaps remain relating to a lack of data, particularly for NSD projects.

13 HYDRO ENERGY↗

Improved production and purification of 240Am by deuteron-induced activation of 240Pu target

An optimized method for production of 240Am via the 240Pu(d,2n)240Am reaction is reported. The optimized method produced 7.94 × 108 ± 7% atoms 240Am/mg 240Pu/µA-h. The yield and purity of the products from the optimized method are evaluated in context of the intended application to produce material to measure the cross-section for the 240Am(n,f) reaction. The presented method is also compared with other production methods available in the literature. Keywords—240Am production, deuteron-induced reactions, cross-section, target preparation, post-irradiation purification Abbreviations GEA—Gamma Energy Analysis

Morrison, Erin C.↗

Depletion of atmospheric neutrino fluxes from parton energy loss

The phenomenon of fully coherent energy loss (FCEL) in the collisions of protons on light ions affects the physics of cosmic ray air showers. As an illustration, we address two closely related observables: hadron production in forthcoming proton-oxygen collisions at the LHC, and the atmospheric neutrino fluxes induced by the semileptonic decays of hadrons produced in proton-air collisions. In both cases, a significant nuclear suppression due to FCEL is predicted. The conventional and prompt neutrino fluxes are suppressed by ~10...25% in their relevant neutrino energy ranges. Previous estimates of atmospheric neutrino fluxes should be scaled down accordingly to account for FCEL. Keywords: Atmospheric neutrinos, Parton energy loss

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PDFDataExtractor: A Tool for Reading Scientific Text and Interpreting Metadata from the Typeset Literature in the Portable Document Format

The layout of portable document format (PDF) files is constant to any screen, and the metadata therein are latent, compared to mark-up languages such as HTML and XML. No semantic tags are usually provided, and a PDF file is not designed to be edited or its data interpreted by software. However, data held in PDF files need to be extracted in order to comply with opensource data requirements that are now government-regulated. In the chemical domain, related chemical and property data also need to be found, and their correlations need to be exploited to enable data science in areas such as data-driven materials discovery. Such relationships may be realized using text-mining software such as the “chemistry-aware” natural-language-processing tool, ChemDataExtractor; however, this tool has limited data-extraction capabilities from PDF files. This study presents the PDFDataExtractor tool, which can act as a plug-in to ChemDataExtractor. It outperforms other PDF-extraction tools for the chemical literature by coupling its functionalities to the chemical-named entityrecognition capabilities of ChemDataExtractor. The intrinsic PDF-reading abilities of ChemDataExtractor are much improved. The system features a template-based architecture. This enables semantic information to be extracted from the PDF files of scientific articles in order to reconstruct the logical structure of articles. While other existing PDF-extracting tools focus on quantity mining, this template-based system is more focused on quality mining on different layouts. PDFDataExtractor outputs information in JSON and plain text, including the metadata of a PDF file, such as paper title, authors, affiliation, email, abstract, keywords, journal, year, document object identifier (DOI), reference, and issue number. With a self-created evaluation article set, PDFDataExtractor achieved promising precision for all key assessed metadata areas of the document text.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Systems Engineering and Analysis in Support of a US Federal Staging Facility for UNF

The US Department of Energy Office of Nuclear Energy (DOE-NE) Office of Spent Fuel and High-Level Waste Disposition is examining a set of system options and conducting supporting analyses to inform the development of an integrated waste management system, which may include one or more federal staging facilities (FSFs) for used nuclear fuel (UNF ) sited using a collaborative siting process. This paper focuses on the ongoing activities in two systems engineering and analysis work areas: (1) data and tools development, validation, and maintenance and (2) systems engineering execution. Within the first work area, the STANDARDS 5.0 UNF data and analysis tool, formerly known as UNF-ST&DARDS, is being developed as a foundational resource to assist in the management of UNF data. It has the key capability to model UNF throughout the entire back end of the fuel cycle. STANDARDS also includes several compatible analysis tools for the time-dependent characterization of UNF and related systems by interfacing with the SCALE code system for nuclear analysis and COBRA-SFS for thermal analysis. Also, within the data and tools area is the Next Generation System Analysis Model (NGSAM), which is an agent-based simulation software tool expressly designed to be capable of modeling the waste management system, including the transportation of UNF to and from a FSF. NGSAM has been developed to enable informed decision-making by providing the capability to analyze various potential system options for the management of UNF and high-level radioactive waste. Finally, in the systems engineering execution area, the team has begun to apply a disciplined systems engineering approach at the system level along with supporting analysis to guide the development of the FSF project requirements (including associated transportation infrastructure). Systems engineering principles and practices and their adaptation/application to design and development activities will ensure that the waste management system is effectively implemented as work proceeds. Other activities include investigating the implications of changes in various assumptions and parameters related to waste management systems, such as UNF acceptance rates, receipt logic, facility capacities and capabilities, use of standardized canisters, and different assumed facility operation start dates. Keywords: federal staging facility (FSF), used nuclear fuel (UNF), integrated waste management (IWM) system, Next Generation System Analysis Model (NGSAM), STANDARDS, systems engineering

Joseph, Robert↗

Bibliometric review and recent advances in total scattering pair distribution function analysis: 21 years in retrospect

Global research activities have been driven by the quest to develop and characterize novel materials for technological advancements. The total scattering pair distribution function (TSPDF) is a powerful and versatile characterization technique for examining the structural details of diverse complex materials including liquid, amorphous, disordered crystalline, and nanostructured materials. Thus, it is critical to keep track of research progress, identify research gaps, and future research directions of the application of the TSPDF technique in materials development and discovery. In this work, a bibliometric analysis of literature regarding the TSPDF technique between 2000 and 2021 was conducted using datasets retrieved from the Web of Science database. The research trends based on publication outputs, research subject distribution, co-authorships among institutions, countries/regions, co-citation of referenced sources, and keyword co-occurrence are evaluated and discussed herein. The impact of the TSPDF technique is projected to increase due to its importance in probing emerging functional materials, and the advances in specialized facilities and instrumentation among the scientific communities engaged with it. Finally, current and emerging research hotspots related to TSPDF technique such as catalysis, computer modeling and simulation, pharmaceutics, machine learning, hydrogen storage, battery materials, and layered structured materials are also identified and discussed.

36 MATERIALS SCIENCE↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗