Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Software quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria↗

Serpentine Magnet Designs for the Interaction Region of the Electron-Ion Collider (EIC)

The Electron-Ion Collider (EIC), hosted by Brookhaven National Laboratory, is designed to deliver a peak luminosity of 1 × 10 34 cm −2 sec −1 . The interaction region (IR) of the EIC imposes several constraints in terms of field quality, aperture, and spatial layout, which necessitates the development of several unique superconducting serpentine direct wind magnets. These magnets are constructed using either a single strand or a small-diameter 6-around-1 NbTi cable, presenting unique challenges for design and optimization. This paper introduces a new computational code specifically developed to streamline and integrate the design process for these magnets, enabling faster design iterations while addressing their complex requirements. Here, in this paper, we first introduce the code, which builds on established electromagnetic fundamentals. The code incorporates tools for optimizing winding patterns and for correcting magnetic multipoles; additionally, it interfaces with established magnet design software. We also present the design of several serpentine magnets for the EIC IR, demonstrating the code’s capability to deliver precise and efficient solutions. These designs highlight the code’s ability to accelerate the development cycle, ensuring the serpentine magnets meet the demanding specifications of the EIC project.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Commissioning of the Mu2e tracker DAQ, planning for the Vertical Slice Test and pre-pattern recognition studies

The primary objective of the Mu2e experiment at Fermilab is to search for the neutrino-less coherent $\mu \rightarrow e$ conversion in the field of an aluminum nucleus ($\mu^- \text{Al} \rightarrow e^- \text{Al}$). The signature of this process is a monochromatic Conversion Electron (CE) with an energy of approximately 104.97 MeV \cite{bartoszek2015mu2e}. Within the Standard Model (SM), the branching ratio for this process, including neutrino masses and oscillation, is expected to be less than $\mathcal{O}(10^{-50})$. This value is far beyond current experimental capabilities. However, models of physics beyond the SM predict much higher relative rates, approaching an observable level. The SINDRUM II experiment set an upper limit on muon conversion at $7 \times 10^{-13}$ (90\% CL) on Au target \cite{SINDRUMII:2006dvw}, and the Mu2e collaboration aims to improve this limit by four orders of magnitude. Observing this process would provide a clear evidence of physics beyond the Standard Model. A brief discussion of the theoretical and experimental aspects is provided in Chapter \ref{intr}. Mu2e adopts a sophisticated experimental setup to achieve its goals, further described in Chapter \ref{mu2echapter}. The central part of the Mu2e detector is the tracker, that consists of 18 tracking stations. The tracker must provide excellent momentum resolution, approximately 1 MeV/c, to distinguish the monochromatic CE signal from the background. To minimize the energy losses, a straw tube tracker will be used \cite{bobbb}. Chapter \ref{chaptertrk} provides an overview of the straw tracker design and its working principles. This Thesis presents a comprehensive study of the Mu2e tracker, covering complementary aspects from initial commissioning to optimization and first steps of the calibration processes. My work at Fermilab has been focused on the complete Data Acquisition (DAQ) testing from both hardware and software perspectives. I was involved in the commissioning of the Mu2e DAQ system and the Vertical Slice Test (VST) of the tracker. The VST encompasses the entire testing chain, from the straws to the readout, and to processed data on disk. I was also focused on the offline analysis, especially on pre-pattern recognition studies, to explore the best methods for identifying $\delta$-electrons during the data taking. Chapter \ref{commissioning} details the commissioning of the tracker DAQ system, emphasizing the importance of understanding of the readout process before the data acquisition. This includes validating the readout logic and firmware through Monte Carlo simulations to confirm functionality and buffering, monitoring the quality of the data from the tracker preamplifiers and front-end electronics, and assessing overall DAQ performance to ensure reliability during future calibration and data-taking. Chapter \ref{planning} discusses the initial steps towards the tracker calibration. The ultimate goal is to perform a time calibration of the first assembled station of the tracker using cosmic muons, aiming for a longitudinal hit position resolution better than 4 cm. This involves determining the signal propagation times and channel-to-channel delays. I performed a Monte Carlo study to determine the impact of the station orientation on the quality of the calibration, in particular on the cosmic track reconstruction, focusing on potential biases that could arise. These studies provide essential insights into the operation, optimization, and calibration of the Mu2e tracker system. Given the high data volume expected during Mu2e operations, estimated at approximately 7 PBytes per year, optimizing memory usage and minimizing CPU consumption are critical. A significant challenge lies in effectively flagging $\delta$-electron hits, which are the primary source of hits in the tracker, without compromising the efficiency of CE hit detection and track reconstruction. A detailed study of pre-pattern recognition and a thorough comparison of two $\delta$-electron flagging algorithms is provided in Chapter \ref{delta}. In Chapter \ref{conclusions}, the findings are concisely summarized, offering a comprehensive synthesis of the research and emphasizing the key insights derived from this study.

43 PARTICLE ACCELERATORS↗

R&D to Ensure a Scientific Basis for Qualification Tests and Standards (Final Report)

Project return on investment in a photovoltaic (PV) system depends increasingly on maintaining high energy yields, and the system lifetime is a major factor in levelized cost of electricity (LCOE). Thus, the rate of PV deployment and the success of these assets depends upon reliable long-term power generation. The overarching objective of this program is to improve photovoltaic (PV) module reliability via development of tests and standards. Where reliability problems or risk are discovered, we can design tests to ensure that these liabilities don't affect future generations of products. Customers can use these tests to understand which products are susceptible to certain degradation mechanisms, and manufacturers can use the tests to design unwanted characteristics out of their products. The work under this program identifies PV reliability needs, performs characterization that provides scientific understanding of targeted degradation mechanisms, and translates those data into practical and predictive test protocols and standards. Major accomplishments include: A model for polarization-type potential induced degradation (PID-p) was developed and validated against experimental data. NREL is currently leading a new edition of IEC 62804-1 for PID detection. PID-p can cause large losses in current and voltage for some module designs on cloudy days. Finite element modeling (FEM) and experiment was used to determine when cells crack in a module. It was shown that cells in landscape orientation are much more likely to crack than those on portrait orientation. Shortly thereafter, the first products with portrait-oriented cells were introduced. Studies of how to test for light and elevated temperature degradation (LeTID) culminated with the publication of IEC TS 63342. Software to predict the progression of LeTID was developed, validated, and made publicly available. Field validated tests and international standards for durability of PV module coatings abrasion, backsheets, and encapsulants were developed. Examples are IEC 62788-1-1, IEC 62788-2 ED2, IEC TS 62788-7-2, IEC 62788-7-3 ED1, IEC 63209-2. NREL led the development a high-temperature testing technical specification, and published guidelines that enable installers to determine whether higher-temperature testing is needed, simply based on location and mounting configuration. In a number of our case studies, variations in the bills of materials or workmanship have been associated with variations in reliability. These observations emphasize the importance of quality assurance to reliability. A framework for criticality (i.e. Pareto) analysis was developed and published. The framework helps us and other researchers determine what problems should be addressed for reliability research to have the biggest industry impact. NREL continues to participate actively in international standards development and stakeholder engagement activities, including organizing an annual PV Reliability Workshop. These activities are important for ensuring we address issues that are relevant and timely, and that we convey our results to those who may benefit.

14 SOLAR ENERGY↗

A product data network to enable faster, easier, and better planning of building envelopes

The building envelopes contributes significantly to the energy-efficiency of the building. Building performance simulation has made it possible to compare façade technologies regarding energy demand, daylighting, thermal and visual comfort in detail. Planners, such as architects and engineers, need experience to find product data with the right quality and level of detail, and to process the data to fit the calculation and the application. In the available time, planners can compare only a limited number of products, which means that better solutions could go unnoticed. This paper presents a new concept for making product data easily accessible for building façade planning. The concept consists of a network of databases for the efficient exchange and use of optical and calorimetric data of glazing units, shading devices, and combinations of both. The paper presents the research questions, an analysis of the current challenges, six design goals for the product data network and its implementation together with a discussion. Many product data sources can be connected to many planning software applications via the specified application programming interface. When planning software connects to the product data network, the planning of building envelopes can be much faster because planners do not need to spend so much time to search and process product data manually. The planning of building envelopes can also become much easier, especially for planners with limited experience. They do not need to understand all the details about which data fits which calculation if the software company implements this. The planning of building envelopes can become much more reliable when software companies validate their use of the product data network, because the current manual process is prone to errors. The planning of building envelopes can also improve because more products can be compared in the available time, allowing better solutions to be found.

Maurer, Christoph↗

Demonstrating Advanced Sensors for In-Situ Monitoring Towards Qualification of Nuclear Relevant Components

The U.S. Department of Energy’s Office of Nuclear Energy Advanced Materials and Manufacturing Technologies (AMMT) program is pursuing qualification of laser powder bed fusion (LPBF) components for nuclear applications. A major focus of this effort is the use of in situ process monitoring and machine learning–based tools to establish real-time quality assurance. The primary objective of this report is to identify and evaluate the most relevant in situ sensor systems for LPBF, and to document the deployment of these systems across platforms critical to the AMMT program. This work demonstrates how in situ monitoring can detect process anomalies, track geometry-dependent flaws, and identify limiting combinations of processing parameters—particularly those related to energy density and complex geometries (e.g., overhanging structures). To support this goal, a diverse suite of sensor modalities was evaluated across LPBF platforms, including visible and near-infrared (NIR) imaging, fringe projection profilometry, long-wavelength infrared (LWIR) thermography, and high-speed photodiode/pyrometry systems. These sensor streams were integrated with Peregrine, a machine-agnostic software platform that, among other capabilities, can generate real-time process anomaly classification. This report documents sensor deployments on multiple AMMT flagship platforms, including the Concept Laser M2 and Renishaw AM400/AM250 systems. Calibration builds with complex, flaw-prone geometries such as unsupported overhangs, stepped features, and thin walls, were used to evaluate how well Peregrine and its associated sensors could detect process anomalies and other instabilities under varied energy densities. It will be shown how Peregrine reliably identifies common process anomalies such as recoater streaking, superelevation, etc., and can be used in post-build analysis for anomaly spatial distributions throughout the build height to better understand the impact of geometry and processing parameter choice on the build. This work demonstrates measurable progress toward the vision that components can be born-qualified by establishing a real-time monitoring framework, identifying limiting process conditions, and laying the foundation for sensor fusion–enabled prediction pipelines that are scalable across platforms and applicable to nuclear-relevant components.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predictive Chemical Kinetic Modeling: Where We Succeed, Where We Struggle, and What Comes Next

Chemical kinetic modeling plays a foundational role in fields ranging from energy to environmental science, pharmaceuticals, and advanced materials. The past two decades have seen remarkable progress, particularly in modeling gas-phase reactions for thermochemical processes, leading to impactful industrial applications such as steam cracking and air quality management. However, new challenges are emerging. The successful development of systematic methodologies for the description of gas-phase kinetics opens the possibility to apply the same approach to the study of more challenging systems. Here, we review recent advances, including ab initio transition state theory-based master equation estimation of elementary rates, automated mechanism generation, machine-learning-assisted kinetics, and uncertainty quantification, and discuss the advances needed to apply the same methodological approach in areas such as heterogeneous catalysis, electrochemistry, liquid-phase and solid-state reactivity, and multiscale model integration. We advocate for the development of targeted tools, especially methods that go beyond empirical tuning toward first-principles-based predictions. We highlight the need for accessible software and AIaugmented workflows to democratize modeling for industry and academia alike. In this perspective, we call attention to not only what has worked but also what remains unsolved, advocating to avoid overemphasizing successes in scientific works at the expense of realism. The next decade should focus on predictive capability, physical accuracy, and community infrastructure (e.g., databases and services) to enable innovation across diverse fields. We argue that kinetic modeling, properly equipped, can accelerate discovery far beyond its traditional domains.

ab initio calculations↗

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron↗

Coupling of high-resolution mass spectrometer and photosynthesis system for comprehensive leaf volatile metabolite profiling

Background Leaf-level biogenic volatile organic compounds (BVOCs) emissions represent a major source of organic gases in the atmosphere, influencing both climate and air quality. These emissions are strongly driven by environmental perturbations, which affect individual plant- to ecosystem-level processes. Uncovering all the BVOCs and understanding how their emissions respond to altered environmental conditions provide critical insights into vegetation-driven changes in atmospheric chemistry. We developed a tandem instrumentation setup that integrates a proton transfer reaction time-of-flight mass spectrometer (PTR-ToF-MS) with parts-per-trillion detection limits and a photosynthetic infrared gas exchange system for the untargeted survey of all the BVOCs. This novel system enables simultaneous, real-time monitoring of BVOC emissions and photosynthetic parameters at the leaf level, offering new opportunities to disentangle the physiological and environmental drivers of VOC release. Furthermore, we established the VOC Analysis and Processing Optimization Resource (VAPOR), an open-access software tool designed for rapid data post-processing and the analysis of the variability of hundreds of BVOCs. We assessed the performance of the tandem system under varying background conditions, using standard gas mixtures and a range of environmental factors. Results Blank emissions were substantially lower for major BVOCs (e.g., isoprene) compared to those observed in plant emissions. Despite this, the observation of background-level VOCs highlights the importance of routinely acquiring and accounting for blank measurements in analyses using the coupled instrumentation. Introduction of known VOC concentrations to the system demonstrated a linear response across different compounds with varying molecular compositions, indicating minimal gas loss regardless of chemical moieties within the coupled instrumentation. We applied the optimized system to investigate the physiological mechanisms driving BVOC emissions across different genotypes of poplar and pennycress. The high mass resolution capabilities of the PTR-ToF-MS, coupled with comprehensive VAPOR-driven data analysis, enabled the identification of several important BVOCs, including methanol and methanethiol; these BVOCs displayed substantial variation across pennycress genotypes and showed concentrations ~ 100–350% higher than the blank. Moreover, isoprene emissions varied significantly among poplar genotypes grown in different potting media. Conclusions Tandem instrumentation offers a powerful tool for profiling volatile molecular markers and elucidating their genetic and environmental underpinnings. This approach enhances our ability to predict BVOC emissions in response to genotype by environmental interactions and contributes to a deeper understanding of vegetation responses to environmental changes.

Biogenic volatile organic compounds↗

Better Climate Challenge Working Groups Non-Energy Benefits of Energy Projects-Improving Financial Payback

Energy efficiency is a key strategy recently identified by the United States Department of Energy as a pillar of industrial decarbonization. For manufacturing companies, improving energy efficiency will reduce money spent on energy utilities such as gas, electricity, and oil. Energy improvement projects also provide valuable benefits outside of simple operating cost reductions, such as reducing the carbon footprint, improving safety metrics and even enhancing quality and productivity. Unfortunately, energy efficiency projects have typically faced an adoption gap, even when they meet criteria such as payback period for capital projects. The inclusion and quantification of non-energy benefits (NEBs), also known as co-benefits, in the decision-making process for energy efficiency projects can improve the overall financial payback periods for those projects as well as potentially improve the company's key performance metrics aligned with business strategies. There are no readily available tools that facilitate this, however, and the most used tools for energy audits address NEBs in a perfunctory way if at all. We integrated research for finding and quantifying non-energy benefits of energy efficiency projects into a commonly recognized continuous improvement practice, the Define, Measure, Analyze, Improve and Control (DMAIC) Process. This process, along with software and supplemental materials, guides energy assessments to find and to quantify NEBs associated with energy conservation opportunities. Our aim is to deliver an easy to use and effective process and software tool and to maximize return on investment for energy efficiency projects as well as contribute to companies' strategic performance goals.

DMAIC↗

A general Bayesian algorithm for the autonomous alignment of beamlines

Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.

47 OTHER INSTRUMENTATION↗

Data for Roebuck et al. (2025), "Differences in dissolved organic matter composition between rivers and estuaries is conserved across freshwater and saltwater coastal regions"

Dissolved organic matter (DOM) in coastal surface waters influences local water quality and is an important component of biogeochemical cycling in coastal systems, but the processes that alter DOM composition along lower reaches of rivers and estuarine waters are poorly understood. Roebuck et al. (2025) leveraged a spatially distributed community sampling effort in coastal ecosystems across two regions to identify broad spatial drivers of surface water DOM composition and identify transferable trends between saltwater and freshwater coastal systems. Samples were collected by community members from 47 locations within the mid-Atlantic and Great Lakes coastal regions.This dataset includes:* A selection of commonly reported absorbance and fluorescence peaks normalized to dissolved organic carbon concentrations* Parallel factor output from the EC1 fluorescence datasets* A selection of commonly reported absorbance and fluorescence peaks * Spectral indices output from matlab script for absorbance and fluorescence datasets* CO2sys calculations of pH changes under varying temperatures and a constant salinity, DIC, and alkalinity concentrationAll data files are plain-text CSV (comma separated value) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

Geomagnetically Induced Current Field Test on Large Grid-Connected Power Transformers: Analysis, Model Development, and Simulations

Geomagnetic-induced current (GIC) flow in power grids can cause undesirable effects such as transformer overheating, harmonics, higher reactive power demand, etc. Many simulation models have been developed to study these effects, but real-world verification on modern transformer designs is rare. Here, this paper presents the first long-duration GIC field test in the U.S. performed on high-voltage, grid-connected transformers featuring winding clamps and tie rods instead of conventional tie bars. Field measurements were taken to evaluate GIC effects. These measurements also aided in developing and validating thermal and electromagnetic transient (EMT) models of the transformers. During the test, significant current and voltage distortions were observed along with considerable transformer reactive power losses. Analysis of the field measurements showed that the transformers’ hottest spot was at the inner windings, and their k-factors were close to factory test and software default values. Thermal simulations indicated that the transformers would not violate their thermal limits even for a GIC waveform that peaks at about 200 A/phase. EMT simulations revealed that increased transformer loading may reduce GIC-induced reactive power demand and harmonics in certain scenarios. The study also highlighted potential inaccuracies in using the k-factor method to calculate transformer reactive power losses.

EMTDC↗

Multi-Task with Procter and Gamble (CRADA No. NFE-10-02672)

The purpose of this Cooperative Research and Development Agreement (CRADA) between UT-Battelle, LLC (the “Contractor) and Procter & Gamble Company (the “Participant”) is the development of a research partnership to create new tools, tests and analytical methods to improve the performance, safety and/or environmental quality of chemicals, advanced materials, food products and manufacturing processes. The Participant operates in three global business units: Beauty, Health and Well-Being and Household Care. Some of its worldwide products include Head and Shoulders®, Pantene®, Gillette® razors and personal care products, Crest®, Dawn®, Tide®, Bounty®, Duracell® batteries; and Iams® pet food among others. At its core, however, the Participant is a science driven company. It supports one of the most robust industrial research and development (R&D) programs in the world. The Participant uses this rich foundation of science to drive innovation across all of its product lines. But the innovation process is not confined in-house The Participant pursues an “open innovation” policy, seeking partnerships with scientists and researchers in universities and national laboratories where it can contribute its extensive knowledge assets and collaborate to advance scientific understanding. The research under this multi-task CRADA was directed under the following general task areas and, throughout the duration of this CRADA the work statement was modified to match the needs of the Parties and the direction of the research. (1) Software modeling, simulation and development; (2) Manufacturing Technologies; (3) Supply Chain Optimization, (4) Advanced Materials.

36 MATERIALS SCIENCE↗

guppy i : a code for reducing the storage requirements of cosmological simulations

ABSTRACT As cosmological simulations have grown in size, the permanent storage requirements of their particle data have also grown. Even modest simulations present a major logistical challenge for the groups which run these boxes and researchers without access to high performance computing facilities often need to restrict their analysis to lower quality data. In this paper, we present guppy, a compression algorithm and code base tailored to reduce the sizes of dark matter-only cosmological simulations by approximately an order of magnitude. guppy is a ‘lossy’ algorithm, meaning that it injects a small amount of controlled and uncorrelated noise into particle properties. We perform extensive tests on the impact that this noise has on the internal structure of dark matter haloes, and identify conservative accuracy limits which ensure that compression has no practical impact on single-snapshot halo properties, profiles, and abundances. We also release functional prototype libraries in C, Python, and Go for reading and creating guppy data.

79 ASTRONOMY AND ASTROPHYSICS↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

Moving toward automated µFTIR spectra matching for microplastic identification: addressing false identifications and improving accuracy

Abstract Infrared spectroscopy is a widely used tool for studying microplastics and identifying microparticles. Researchers rely on spectral libraries to differentiate between synthetic and natural materials. Unfortunately, spectral library matching is not perfect, and best practices require researchers to use time consuming, manual peak matching to assess spectral matches. Moving toward automated matching requires increased confidence in the matching process. Using spectra matching software may increase the efficiency of particle identification, however some matching strategies may confuse natural materials such as cotton, silk, and plant matter with common classes of synthetics such as polyesters and polyamides. In this experiment, we prepared 22 pristine sample materials from natural and synthetic sources and measured micro-Fourier transform infrared (µFTIR) spectra in transmission mode for each sample using a Thermo Nicolet iN10 MX instrument. The collected spectra were then input into two spectral library matching systems (Omnic Picta and Open Specy), using a total of five identification routines. Next, we placed a subset of four pristine microplastic materials in a biologically active river system for two weeks to simulate environmental samples. These simulated environmental samples were processed using 10% hydrogen peroxide for 24 h to remove organic contamination and then identified using the strongest performing library. We found that libraries with fewer sample spectra produced lower correlation matches and that using derivative correction greatly reduced the number of inaccuracies in identifying materials as either natural or synthetic. We also found that environmental fouling reduced the correlation value of library matches when compared to pristine particles, however the effect was not consistent across the four materials tested. Overall, we found that the accuracy of automated library matching in the tested systems and processing routines varied from 64.1 to 98.0% for distinguishing between natural and synthetic materials, and that a high Hit Quality Index (HQI) did not always correlate with accuracy. These results are important for the microplastic field, demonstrating a need to rigorously test spectral libraries and processing routines with known materials to ensure identification accuracy.

Kozloski, Rachel↗

Using a Large Language Model as a Building Block to Generate Usable Validation and Verification Suite for OpenMP

In the HPC area, both hardware and software move quickly. Often new hardware is developed and deployed, the corresponding software stack, including compilers and other tools, are under active development while leading edge software developers are working to port and tune their applications, all at the same time. While the software ecosystem is in flux, one of the key challenges for users is obtaining insight into the state of implementation of key features in the programming languages and models their applications are using – whether they have been implemented, and whether the implementation conforms to the specification, especially for newly implemented features (less tested by widespread use). OpenMP is one of the most prominent shared memory programming models used for on-node programming in HPC. With the shift towards accelerators (such as GPUs and FPGAs) and heterogeneous programming OpenMP features are getting more complex. It is natural to ask whether generative AI approaches, and large language models (LLMs) in particular, can help in producing validation and verification test suites to allow users better and faster insights into the availability and correctness of OpenMP features of interest. In this work, we explore the use of ChatGPT-4 to generate a suite of tests for OpenMP features. We have chosen a set of directives and clauses, a total of 78 combinations, which first appeared in OpenMP 3.0 (released in May 2008) but are also relevant for accelerators. We prompted ChatGPT to generate tests in the C and Fortran languages, for both host (CPU) and device (accelerator). On the Summit super-computer using the GNU implementation, we found that, of the 78 generated tests 67 C tests and 43 Fortran tests compiled successfully and fewer than those executed to completion. On further analysis we show that not all generated tests are valid. We document the process, results, and provide detailed analysis regarding the quality of tests generated. With the aim of providing input to a production quality validation and verification suite, we manually implement the corrections required to make the tests valid according to the current OpenMP specification. We quantify this effort as small, medium, or large, and record the lines of code changed to correct the invalid tests. With the corrected tests we validate recent implementations from HPE, AMD, and GNU on the Frontier supercomputer. Our experiment and subsequent analysis show that although LLMs are capable of producing HPC specific codes, they are limited by their understanding of the deeper semantics and restrictions of programming models such as OpenMP. Unsurprisingly more commonly used features have better support, while some OpenMP 3.0 directives such as sections and tasking are not universally supported on accelerators. We demonstrate that successful compilation and execution to completion are inadequate metrics for evaluating generated code and that, at this time, commodity LLMs require expert intervention for code verification. This points to gaps in the training data that is currently available for HPC. We demonstrate that with "small" effort 37% of generated invalid C tests and 63% of generated invalid Fortran tests could be corrected. This improves productivity of test generation as we circumvent writing from scratch and the common programming errors associated with it.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗