Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Heterogenous electromediated depolymerization of highly crystalline polyoxymethylene

Abstract Post-consumer plastic waste in the environment has driven the scientific community to develop deconstruction methods that yield valued substances from these synthetic macromolecules. Electrocatalysis is a well-established method for achieving challenging transformations in small molecule synthesis. Here we present the first electro-chemical depolymerization of polyoxymethylene—a highly crystalline engineering thermoplastic (Delrin®)—into its repolymerizable monomer, formaldehyde/1,3,5-trioxane, under ambient conditions. We investigate this electrochemical deconstruction by employing solvent screening, cyclic voltammetry, divided cell studies, electrolysis with redox mediators, small molecule model studies, and control experiments. Our findings determine that the reaction proceeds via a heterogeneous electro-mediated acid depolymerization mechanism. The bifunctional role of the co-solvent 1,1,1,3,3,3-hexafluoro-2-propanol (HFIP) is also revealed. This study demonstrates the potential of electromediated depolymerization serving as an important role in sustainable chemistry by merging the concepts of renewable energy and circular plastic economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

The Zap Energy approach to commercial fusion

Zap Energy is a private fusion energy company developing the sheared-flow-stabilized (SFS) Z-pinch concept for commercial energy production. Spun out from the University of Washington, these experimental and computational efforts have resulted in devices with quasi-steady DD fusion yields above 109 per pulse. These devices support scaling toward energy breakeven on existing devices as well as beyond to commercially relevant engineering fusion gains. This article discusses the strategy behind Zap's development path, which is derived directly from the engineering and scientific elegance of the confinement method. Without need for external confinement or heating technologies, the SFS Z pinch relies on plasma self-organization. This compact magnetic confinement technology could, in turn, provide the basis for a cost-effective fusion power plant, vastly reduced in complexity from its competitors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Characterization of D-T Generator Neutron Flux Spectrum for Cyclic Neutron Activation Analysis Experiments

Improving nuclear data for short-lived fission product yields will further our fundamental understanding of fission, which is needed across various scientific fields and applications. One method of attaining the needed product yield data is through cyclic neutron activation, which allows a target to be irradiated in a neutron environment and then transported for counting of the radionuclides produced, typically via g spectroscopy. Recently, such a system has been constructed and commissioned at Pacific Northwest National Laboratory. Targets are shuttled between the head of a D-T neutron generator and a counting station with a transit time of 2 seconds. As part of the characterization of this system, the neutron flux was studied using two activation targets. The neutron flux from the deuterium-tritium fusion generator was determined to be 9.95(33) x10^8 n/cm2-s with a peak energy of 14.9 MeV and a spread of approximately 4 decades between the epithermal and 14 MeV peak group flux.

cyclic neutron activation, gamma-ray spectroscopy,↗

Toward a Holistic Performance Evaluation of Large Language Models Across Diverse AI Accelerators

Artificial intelligence (AI) methods have become critical in scientific applications to help accelerate scientific discovery. Large language models (LLMs) are being considered a promising approach to address some challenging problems because of their superior generalization capabilities across domains. The effectiveness of the models and the accuracy of the applications are contingent upon their efficient execution on the underlying hardware infrastructure. Specialized Al accelerator hardware systems have recently become available for accelerating Al applications. However, the comparative performance of these AI accelerators on large language models has not been previously studied. In this paper, we systematically study LLMs on multiple AI accelerators and GPUs and evaluate their performance characteristics for these models. We evaluate these systems with (i) a micro-benchmark using a core transformer block, (ii) a GPT-2 model, and (iii) an 1,I,M-driven science use case, GenSLM. We present our findings and analyses of the models' performance to better understand the intrinsic capabilities of AI accelerators. Furthermore, our analysis takes into account key factors such as sequence lengths, scaling behavior, and sensitivity to gradient accumulation steps.

Emani, Murali↗

Characterizing Machine Learning I/O Workloads on Leadership Scale HPC Systems

High performance computing (HPC) is no longer solely limited to traditional workloads such as simulation and modeling. With the increase in the popularity of machine learning (ML) and deep learning (DL) technologies, we are observing that an increasing number of HPC users are incorporating ML methods into their workflow and scientific discovery processes, across a wide spectrum of science domains such as biology, earth science, and physics. This gives rise to a diverse set of I/O patterns than the traditional checkpoint/restart-based HPC I/O behavior. The details of the I/O characteristics of such ML I/O workloads have not been studied extensively for large-scale leadership HPC systems. This paper aims to fill that gap by providing an in-depth analysis to gain an understanding of the I/O behavior of ML I/O workloads using darshan - an I/O characterization tool designed for lightweight tracing and profiling. We study the darshan logs of more than 23, 000 HPC ML I/O jobs over a time period of one year running on Summit - the second-fastest supercomputer in the world. This paper provides a systematic I/O characterization of ML I/O jobs running on a leadership scale supercomputer to understand how the I/O behavior differs across science domains and the scale of workloads, and analyze the usage of parallel file system and burst buffer by ML I/O workloads.

Paul, Arnab↗

Estimating Eigenenergies from Quantum Dynamics: A Unified Noise-Resilient Measurement-Driven Approach

Ground state energy estimation in physical, chemical, and materials sciences is one of the most promising applications of quantum computing. In this work, we introduce a new hybrid approach that finds the eigenenergies by collecting real-time measurements and post-processing them using the machinery of dynamic mode decomposition (DMD). From the perspective of quantum dynamics, we establish that our approach can be formally understood as a stable variational method on the function space of observables available from a quantum many-body system. We also provide strong theoretical and numerical evidence that our method converges rapidly even in the presence of a large degree of perturbative noise, and show that the method bears an isomorphism to robust matrix factorization methods developed independently across various scientific communities. Our numerical benchmarks on spin and molecular systems demonstrate an accelerated convergence and a favorable resource reduction over state-of-the-art algorithms. The DMD-centric strategy can systematically mitigate noise and stands out as a leading hybrid quantum-classical eigensolver.

Shen, Yizhi↗

Advanced Method Optimization for Sampling and Analysis Instrumentation

This work presents a generalized approach for analytical method optimization that branches the gap between techniques historically employed and accurate modern optimization techniques suitable for various applications. The novelty of the described strategy is the utilization of multivariate, multiobjective optimization with Karush-Kuhn-Tucker conditions to bound the optimization space to solutions within the physical limitations of instrumentation. Briefly, the basic steps outlined in this paper are to (1) determine the objective(s) that should be maximized or minimized based on the goals of the analytical application, (2) conduct a screening experiment, (3) perform ANOVA to determine the parameters which have a statistically significant effect on the objective, (4) conduct an experiment (e.g., Box-Behnken design) to collect data for fitting the objective equation, and (5) determine the physical constraints of the parameters and solve the Lagrangian to determine the optimal method parameters. A broad approach to optimization target selection allows for robust method tuning to develop improved data sets amenable for chemometrics and machine learning algorithm development. Gas chromatography-mass spectrometry was selected as a use case due to its broad use across scientific fields and time-consuming method development involving numerous parameters. In conclusion, this strategy can reduce the cost of research, improve data quality, and enable the rapid development of new analytical technique.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancing electrochemical impedance analysis through innovations in the distribution of relaxation times method

Electrochemical impedance spectroscopy (EIS) is a key tool across various scientific disciplines, including energy sciences, chemistry, and biology, enabling the analysis of electrochemical systems. However, conventional methods for interpreting EIS data are often complex and model dependent. The distribution of relaxation times (DRT) offers a non-parametric approach that simplifies the interpretation process by providing a timescale interpretation of EIS data. This article provides a comprehensive review of current methods for DRT inversion. Additionally, a survey of practitioners highlights key challenges in the field. Here, the findings underscore the need for standardized DRT analysis and benchmarks, as well as the development of automated analysis tools. These advancements would improve the usability and interpretability of EIS data. Ultimately, implementing these improvements could not only propel the field forward but also expand the application of DRT in scientific research by making it accessible to a broader range of researchers, including those without specialized expertise in programming or statistics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advanced Method Optimization with Categorical and Constrained Continuous Parameters

Traditional approaches to analytical method optimization (e.g., univariate and “guess-and-check”) can be time-consuming, costly, and often fail to identify true optima within the parameter space. Previous work defined and implemented a generalized technique for method optimization for continuous method parameters, but a knowledge gap remains for the incorporation of categorical variables into these advanced method optimization schemes. This work presents and validates a generalized optimization approach that incorporates both continuous and categorical variables while also utilizing a multivariate, multiobjective optimization scheme with Karush–Kuhn–Tucker conditions to bound the optimization space to solutions within the physical limitations of the parameter space. Method optimization from a case study using GC–MS for the analysis of 11 analytical standards with objectives to minimize peak width and maximize peak height resulted in a 3 orders of magnitude improvement in the average peak height and a 2 orders of magnitude improvement in the average peak width compared to the least optimal (but reasonable) instrumental parameters utilized in this study. This approach to optimization allows for a customizable method optimization in which users can include both continuous and categorical variables to achieve objectives specific to their analytical goals. This approach significantly reduces the labor and cost associated with traditional method development approaches and can be applied in a variety of scientific fields across a range of laboratory techniques (e.g., instrument method development, sample preparation, and extraction techniques).

Amorphous materials↗

Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing

The Joint Laboratory on Extreme-Scale Computing (JLESC) was initiated at the same time lossy compression for scientific data became an important topic for the scientific communities. The teams involved in the JLESC played and are still playing an important role in developing the research, techniques, methods, and technologies making lossy compression for scientific data a key tool for scientists and engineers. Here, in this paper, we present the evolution of lossy compression for scientific data from 2015, describing the situation before the JLESC started, the evolution of this discipline in the past 8 years (until 2023) through the prism of the JLESC collaborations on this topic and some of the remaining open research questions.

Compression for AI↗

Universal Workflow Language and Software Enable Geometric Learning and FAIR Scientific Protocol Reporting

Written language and conventional data structures for representing scientific procedures suffer from low process detail, often fail to accurately represent protocols, and lack universality. New strategies for the handling of experimental data are needed to provide viable process information for both humans and machines. In this work, we present the universal workflow language (UWL) and interface (UWLi). UWL is a findable, accessible, interoperable, and reusable (FAIR)-compatible, graph-based data architecture that can capture arbitrary scientific procedures through workflow representation, and UWLi is an accompanying software package for building, manipulating, and interpreting UWL entries. The UWL format was found to be highly effective in identifying deficiencies in the reported process details of high-impact, peer-reviewed scientific journals, and in simulated scenarios, the graph format was shown to be more effective than conventional methods in predictively modeling the outcome of diverse scientific protocols. Implementation of UWL could enable more accurate scientific communication and more impactful process datasets.

14 SOLAR ENERGY↗

Neutronics Calculation Advances at Los Alamos: Manhattan Project to Monte Carlo

The history and advances of neutronics calculations at Los Alamos during the Manhattan Project through the present are reviewed. Substantial improvements to neutron diffusion methods and the invention of both the Monte Carlo neutron transport methods in 1947 and deterministic discrete ordinates Sn in 1953 were all made at Los Alamos just after the Manhattan Project. We briefly summarize early simpler and more approximate neutronics methods and then describe the need to better predict neutronics behavior through consideration of theoretical equations, models and algorithms, experimental measurements, and available computing capabilities and their limitations. This paper briefly covers key advances in deterministic methods during the Manhattan Project. These capabilities, coupled with increasing postwar defense needs and the invention of electronic computing with the Electronic Numeric Integrator and Computer, known as ENIAC, and the Mathematical Analyzer Numerical Integrator and Automatic Computer Model, known as MANIAC, led to the creation of Monte Carlo and deterministic discrete ordinates neutronics transport methods. We note the important role that the scientific comradery between the Los Alamos scientists played in the process. This paper briefly covers the early methods, algorithms, computers, and electronic and women pioneers that enabled Monte Carlo to spread to all areas of science. We focus heavily on these early developments and the subsequent creation of the MCNP® code, advances in its associated nuclear data, and its applications to problems of national defense at Los Alamos.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Crowdsourcing Global Perspectives in Ecology Using Social Media

Transparent, open, and reproducible research is still far from routine, and the full potential of open science has not yet been realized. Crowdsourcing–defined as the usage of a flexible open call to a heterogeneous group of individuals to recruit volunteers for a task –is an emerging scientific model that encourages larger and more outwardly transparent collaborations. While crowdsourcing, particularly through citizen- or community-based science, has been increasing over the last decade in ecological research, it remains infrequently used as a means of generating scientific knowledge in comparison to more traditional approaches. We explored a new implementation of crowdsourcing by using an open call on social media to assess its utility to address fundamental ecological questions. We specifically focused on pervasive challenges in predicting, mitigating, and understanding the consequences of disturbances. In this paper, we briefly review open science concepts and their benefits, and then focus on the new methods we used to generate a scientific publication. We share our approach, lessons learned, and potential pathways forward for expanding open science. Our model is based on the beliefs that social media can be a powerful tool for idea generation and that open collaborative writing processes can enhance scientific outcomes. We structured the project in five phases: (1) draft idea generation, (2) leadership team recruitment and project development, (3) open collaborator recruitment via social media, (4) iterative paper development, and (5) final editing, authorship assignment, and submission by the leadership team. We observed benefits including: facilitating connections between unusual networks of scientists, providing opportunities for early career and underrepresented groups of scientists, and rapid knowledge exchange that generated multidisciplinary ideas. We also identified areas for improvement, highlighting biases in the individuals that self-selected participation and acknowledging remaining barriers to contributing new or incompletely formed ideas into a public document. While shifting scientific paradigms to completely open science is a long-term process, our hope in publishing this work is to encourage others to build upon and improve our efforts in new and creative ways.

54 ENVIRONMENTAL SCIENCES↗

Methods of improving brain dose estimates for internally deposited radionuclides *

The US National Council on Radiation Protection and Measurements (NCRP) convened Scientific Committee 6–12 (SC 6–12) to examine methods for improving dose estimates for brain tissue for internally deposited radionuclides, with emphasis on alpha emitters. This Memorandum summarises the main findings of SC 6–12 described in the recently published NCRP Commentary No. 31, ‘Development of Kinetic and Anatomical Models for Brain Dosimetry for Internally Deposited Radionuclides’. The Commentary examines the extent to which dose estimates for the brain could be improved through increased realism in the biokinetic and dosimetric models currently used in radiation protection and epidemiology. A limitation of most of the current element-specific systemic biokinetic models is the absence of brain as an explicitly identified source region with its unique rate(s) of exchange of the element with blood. The brain is usually included in a large source region called Other that contains all tissues not considered major repositories for the element. In effect, all tissues in Other are assigned a common set of exchange rates with blood. A limitation of current dosimetric models for internal emitters is that activity in the brain is treated as a well-mixed pool, although more sophisticated models allowing consideration of different activity concentrations in different regions of the brain have been proposed. Here case studies for 18 internal emitters indicate that brain dose estimates using current dosimetric models may change substantially (by a factor of 5 or more), or may change only modestly, by addition of a sub-model of the brain in the biokinetic model, with transfer rates based on results of published biokinetic studies and autopsy data for the element of interest. As a starting place for improving brain dose estimates, development of biokinetic models with explicit sub-models of the brain (when sufficient biokinetic data are available) is underway for radionuclides frequently encountered in radiation epidemiology. A longer-term goal is development of coordinated biokinetic and dosimetric models that address the distribution of major radioelements among radiosensitive brain tissues.

61 RADIATION PROTECTION AND DOSIMETRY↗

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Assessment of human nuclear and mitochondrial DNA qPCR assays for quantification accuracy utilizing NIST SRM 2372a

In forensic DNA casework, a highly accurate real-time quantitative polymerase chain reaction (qPCR) assay is recommended per the Scientific Working Group on DNA Analysis Methods (SWGDAM) (SWGDAM Validation Guidelines for DNA Analysis Methods [1]) to determine whether a DNA sample is of sufficient quantity and robust quality to move forward with downstream short tandem repeats (STR) or sequencing analyses. Most of these assays rely on a standard curve, referred to herein and traditionally as absolute qPCR, in which an unknown is compared, relative to that curve. However, one fundamental issue with absolute qPCR is the quantifiable concentration of commercial assay standards can vary depending on (1) origin, i.e., whether from a cell line or a human subject, (2) supplier, (3) lot number, (4) shipping method, etc. In 2018, the National Institute for Standards and Technology (NIST) released a human DNA standard reference material for evaluating qPCR quantification standards, Standard Reference Material (SRM) 2372a, Romsos et al. (2018) [2] which contains three well-characterized human genomic DNA samples: Component A) a single male1 donor, Component B) a single female 1 donor, and Component C) a 1:3 male 2 :female 2 donor, each with certification data for nDNA and informational mitochondrial DNA(mtDNA)/nuclear DNA (nDNA) ratio data. The SRM 2372a was used to assess four qPCR assays: (1) Quantifiler Trio (Thermo Fisher Scientific, Waltham, MA) for nDNA quantification, (2) NovaQUANT (EMD Millipore Corporation, San Diego, CA) for nDNA and mtDNA quantification, (3) a custom duplex mtDNA assay, and (4) a custom triplex mtDNA assay. Additionally, extracts from eighteen (18) skeletal remains were tested with the latter three assays for concordance of DNA concentration and with assays (2) and (3), for the degradation state. Our assessment revealed that an accurate, efficient, and reproducible qPCR assay is dependent on (1) the quality and reliability of the DNA standard, (2) the qPCR chemistry, and (3) the specific primers, and probes (if applicable), used in an assay. Finally, our findings indicate qPCR assays may not always quantify as expected and that performance of each lot should be verified using a well-characterized DNA standard such as the NIST SRM 2372a and adjusted if warranted.

59 BASIC BIOLOGICAL SCIENCES↗