Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A radioisotope - enabled reactive transport model for deep vadose zone carbon

In mountainous regions, which constitute the principle source of recharge to major rivers and regional aquifers, infiltration occurs through fractured, partially saturated, weathered bedrock that acts as a boundary layer between saturated aquifers and surface soil. Commonly this deep vadose zone (DVZ) is many meters thick, and yet its role in regulating the generation, retention and mobility of reactive solutes, including nutrients, contaminants and weathering products, is largely unknown. In particular, many of the key reactions that drive the formation of the weathered DVZ and the quality of water moving through it are redox processes, regulated by the availability of organic carbon and oxygen below the soil layer. The role of the DVZ is thus also poorly constrained in the context of carbon stocks and mobility, particularly in lithologies that are naturally high in organic carbon, such as shales. The overarching hypothesis of this study is that upland regions developed in geologic settings with abundant petrogenic carbon store and actively cycle carbon in the weathered DVZ below the soil and above the water table at rates that are significant and currently unconstrained. In order to quantify this cycling, the current study combines novel instrumentation techniques allowing new direct sampling of DVZ systems with advanced numerical reactive transport simulations of carbon transport and transformation. Critically, these simulations will explicitly treat the three isotopes of carbon (the abundant 12C, the stable rare 13C and the radioactive 14C) in a unified framework, thus clearly parsing between the contributions of modern surface derived carbon and lithologic carbon sources in integrated measurements of fluid and gas phase fluxes. This novel model capability will be applied to test the role of DVZ carbon cycling as a regulator of water quality and geological weathering in two complementary field sites both located in organic carbon rich shale lithologies. The first is the Eel River Critical Zone Observatory (ERCZO) in Mendocino County, California, and the second is the Lawrence Berkeley National Laboratory Watershed Function Scientific Focus Area (SFA) in the East River watershed, near Crested Butte, Colorado. At the ERCZO, a novel vadose zone monitoring system has been installed in a 20 m thick, partially saturated, weathered shale hillslope, and preliminary data already indicate substantial CO2 flux generated many meters below the soil surface. At the SFA field site, an instrumented hillslope transect indicates a more complex multi-dimensional fluid and solute transport regime, which will serve as a key test of the calibrated models. Collectively, this project will advance understanding of the cycling of carbon belowground and in relation to transport pathways across the poorly constrained DVZ characteristic of primary water recharge areas. The key product of this work will be enhanced isotope simulation capabilities that are robust and publicly available for application across a broad diversity of systems.

58 GEOSCIENCES↗

Performance Monitoring Program: Developing Comparative Metrics for Fitness-for-Duty Programs

To comply with U.S. Nuclear Regulatory Commission (NRC) regulations, licensees and entities authorized under Title 10 of the Code of Federal Regulations (CFR) Part 26 Section 26.3 (§ 26.3) are required to have fitness-for-duty (FFD) programs. Under § 26.3, the expectation of these FFD programs is to provide reasonable assurance that individuals who are granted unescorted access to nuclear power reactor protected areas and Category I fuel cycle facility material control areas are trustworthy, will perform their tasks in a reliable manner, are not under the influence of any substance, legal or illegal, that may impair their ability to perform their duties, and are not mentally or physically impaired from any cause that can adversely affect their ability to safely and competently perform their duties. Pacific Northwest National Laboratory (PNNL) was tasked with developing a performance monitoring program to risk inform NRC inspection and policy regarding quantitative FFD performance data. To meet this need, PNNL developed methodologies that could be implemented within a performance monitoring program. Throughout this report, these methodologies are referred to as comparative metrics. The NRC provided PNNL with 2016–2019 data from annual reporting forms and single positive test forms provided by licensees and other entities. For most of the comparative metrics, the analyses required customized processing, such as creating filtering fields, adding data fields and summaries, and joining datasets. The comparative metrics developed include: random testing rate, random policy violation rate, pre-access policy violation rate, subversion attempt rate, and number of policy violations by labor category. The comparative metrics can be used for parsing and visualizing FFD program data and monitoring FFD performance at the labor category, facility, licensee, and industry levels to risk inform NRC inspection and policy with regard to the FFD data currently collected from licensees and other entities that implement Part 26 requirements. Furthermore, these developed comparative metrics allow a more in-depth look at the FFD programs for the industry to discern trends and patterns that may warrant changes at an industry level and to inform policy decisions. These analyses should be refreshed as new FFD data become available.

99 GENERAL AND MISCELLANEOUS↗

Report on ISR-1 High-Altitude Balloon Flight

To test small technologies at lower cost for space science applications, LANL has developed a small high altitude balloon payload that could, in the future, be regularly and inexpensively launched from LANL. A neutron detector, NEMO, was integrated to evaluate its performance in a space-like mixed-radiation environment and collect neutron data in the atmosphere. In collaboration with EES-14, a high-altitude balloon payload was launched from LANL Technical Area 51 on February 27, 2023 and April 17, 2023. For real-time geolocation, a SAM-M8Q M8 GNSS module was used to get position and time, and an Iridium RockBLOCK 9603 was used to communicate with the ground using the Iridium satellite fleet. These modules were all controlled using an Iteaduino Mega microcontroller board. Finally, a High Altitude Science Eagle Flight Computer with a temperature pressure sensor ran independently, writing data to an SD card. All of these modules were powered by a 5 mAh lithium polymer battery. The battery was attached to the bottom of the payload while the remaining electronics were embedded in the underside of the top of the payload. These modules were wired as seen in Figure 1-2. The Iteaduino Mega microcontroller board was programmed to use the RockBLOCK to send a message once every 10 minutes containing neutron and GPS data read off the NEMO and SAM M8Q, respectively. Once the message send attempt finished, the RockBLOCK would be slept for the rest of the 10 minute interval. The Eagle Flight Computer ran continuously throughout the flight, taking data every 6 seconds. The RockBLOCK message data was set up to be delivered from the Iridium satellite fleet to a website, where it was stored and parsed to create live maps and plots for analysis and balloon retrieval. The RockBLOCK message data was additionally configured to be sent to an email as a fail-safe. The payload was ground-tested successfully for over 50 hours, with multiple revisions occurring to best prepare for conditions at altitude and improve the software and firmware to fix any issues that cropped up with the data pipeline. Additional to the balloon payload, the flight had an attached iMet-4 radiosonde and Garmin T5 GPS Dog Collar. The radiosonde provided GPS and meteorological data. The T5 dog collar is used along with a Garmin Astro 430 to track the balloon at a range of up to 9 miles for retrieval. The balloon itself was initially a 1600 g meteorological balloon with an attached High Altitude Science parachute, both of which can be seen in Figure 1-3. After the first flight, the EES team swapped to a Rocketman parachute.

42 ENGINEERING↗

Project Update for “Designing Nuclear-data Measurements that Resolve Discrepancies in Existing Data” [Slides]

AIACHNE has made key progress this past year and will contribute to the larger scientific community. We recovered input data for the current 252 Cf(sf) PFNS evaluation that was previously lost. We render a standard to the best of our ability reproducible. We critically reviewed past data as input for ML & new standard evaluation that will impact PFNS of all major actinides. We developed a unique AI/ ML code that highlights which measurement features are related to bias and are working towards open-sourcing it for the community. Features that were identified as related to bias follow physics’ intuition and bring new understanding of exp effects and might help us for other reactions and isotopes. The results highlight that EXFOR is a goldmine of features that could help us understand experiment bias (if they are easy to parse).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bia. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Ref. [2]. The prerequisite for applying machine learning techniques is casting the metadata into a format that can be parsed by the algorithm. This step might seem trivial but requires to find a unique language where metadata that carry the same physics meaning across several experiments must have the same identifier. One example is, for instance, the neutron detector. As seen in Figure 1, the machine learning code identified the use of 6 Li detectors as being related to bias in some datasets of the AIACHNE 252 Cf PFNS experimental database. In fact, here are several experiments that used neutron detectors containing 6Li in the database, for instance for the example below. EXFOR format has a unique keywords describing detectors such as “SCIN” or “GLASD”. One may think that these keywords are already sufficient descriptors for ML to uniquely find an issue. However, “SCIN” (used for [3, 4]) and “GLASD” (used for [5]) fail to inform the algorithm what is the active material in the detector. And, the key common issue leading to bias in 252 Cf related to neutron detectors is not whether it is a glass detector or a scintillator. No, the issue is that 6 Li was within both detector types and that even small mistakes in the detector response functions around approximately 200 keV are amplified by the 6 Li(n,α) resonance there leading to bias in data as highlighted in Fig. 1 and Ref. [1]. Hence, the features describing the neutron detector must call out the active material in the detector, rather than the existing EXFOR detector keyword, that the ML algorithm can find physically meaningful features related to bias. The AIACHNE team used a precursor of the WPEC (Working Party on International Nuclear Data Evaluation Co-operation) SG(Subgroup)-50 format to store the metadata for the ML analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗

Prototyping DAQ network functions on FABRIC

This poster relates to a project that is investigating the use of the FABRIC federated testbed to scale-up research infrastructure to develop network functions that use smart (programmable) network equipment to support High Energy Physics (HEP) research. The poster describes preliminary work that involves non trivial packet parsing related to Fermilab DAQ workloads that is implemented on real hardware, and a correctness and performance evaluation.

Sagstad, Bjoern [IIT, Chicago]↗

Improved Weld Residual Stress Modeling System in BlackBear

This report presents enhancements to the MOOSE-based BlackBear application aimed at improving its capability to simulate welding and other thermo-mechanical manufacturing processes. Two primary avenues of improvement are pursued. First, to enhance user accessibility, we introduce a centralized default block restriction mechanism that ensures coverage checks are performed within user-specified default blocks. This default setting is applied consistently to all block-describable objects, such as variables, kernels, and more. In addition, we develop a modular action for moving heat source simulations, which integrates path file parsing, subdomain modification, and heat source kernel enforcement into a single, streamlined configuration. Second, to improve solver robustness, we implement an alternative method for assigning initial conditions to the updated active domain during the simulation, thereby enhancing convergence behavior. To validate the framework, we design and conduct several benchmark simulations, including heat conduction with progressive material addition, linear elasticity with time-dependent material deposition, and viscoplasticity model with isotropic hardening under similar conditions. Finally, we demonstrate the effectiveness of the proposed framework through large-scale thermo-mechanical welding simulations in both two and three dimensions.

42 ENGINEERING↗

Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing

Due to the growing use of large language models (LLMs) by developers and researchers, it has become essential to reliably evaluate their ability to generate code that uses specialized libraries. We explore the use of compile-time and run-time evaluation of LLM-generated Kokkos code through extending the methods used by OpenAI with the HumanEval dataset. Our evaluation framework is based on the first 40 prompts from the Kokkos138 dataset. We start by discussing two different forms of LLM prompting, using entirely plain English or providing pseudocode for added context. These two methods are used to generate Kokkos code with the Llama-3.1-8B-Instruct and CodeQwen1.5-7B-Chat models. We found that both forms of prompting led to high failure rates and difficulties with reliably parsing LLM-generated code, while prompts with pseudocode for context generally led to improved results on more complicated tests.

97 MATHEMATICS AND COMPUTING↗

Improving Runtime Performance of Tensor Computations using Rust From Python

In this work, we investigate improving the runtime performance of key computational kernels in the Python Tensor Toolbox (pyttb), a package for analyzing tensor data across a wide variety of applications. Recent runtime performance improvements have been demonstrated using Rust, a compiled language, from Python via extension modules leveraging the Python C API—e.g., web applications, data parsing, data validation, etc. Using this same approach, we study the runtime performance of key tensor kernels of increasing complexity, from simple kernels involving sums of products over data accessed through single and nested loops to more advanced tensor multiplication kernels that are key in low-rank tensor decomposition and tensor regression algorithms. In numerical experiments involving synthetically generated tensor data of various sizes and these tensor kernels, we demonstrate consistent improvements in runtime performance when using Rust from Python over 1) using Python alone, 2) using Python and the Numba just-in-time Python compiler (for loop-based kernels), and 3) using the NumPy Python package for scientific computing (for pyttb kernels).

97 MATHEMATICS AND COMPUTING↗

Skipper-CCD Quantum Efficiency Analysis

Scientific skipper-CCDs with single-electron resolution present dozens of possibilities for detecting dark matter candidates. Fermilab's Cosmic Physics Center contributes to the DarkNESS mission, which aims to place a skipper multi-chip module in a 6U CubeSat designed for low earth orbit. The DarkNESS nanosat will have the capability to search for 1-10keV band X-rays that may originate from DM decays. One of the challenges of detection in space is the large amount of cosmic radiation contributing to sensor noise. To mitigate this, an aluminum shield is proposed to be placed on the sensor. This project aims to test and characterize the energy resolution with a shield of various thicknesses (0-100nm) using a single CCD and an iron-55 x-ray source. ROOT analysis was used to parse data from several runs into sections based on shield thickness, create strategic data cuts, and characterize Fano plus signal shot noise in the sensor.

Wells, Megan E. [U. Illinois, Chicago]↗

Remote Instrumentation and Data Acquisition

This poster outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and future work, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Illinois U., Urbana]↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

GridCoPilot for Thermal Events: An LLM-Based Platform for Power Grid Reliability Analysis

Large Language Models show promise for translating natural language into database queries, but deploying such systems in safety-critical domains requires high reliability. We present an application of GridCoPilot to thermal event analysis (heatwaves and coldwaves) that affect power grid reliability. Our approach uses a LangChain SQL Agent to translate natural language queries into auditable SQL statements, with deterministic visualization routines that parse the structured query results. We introduce structural framing as a design principle, we integrate a NERC-region-level event library with county-level meteorology and decompose the combined data into three relational tables (event metadata, county-level event details, and a county-to-NERC subregion mapping), using prompt-guided joins to direct the model toward correct multi-table queries. For two core analytical patterns (identifying worst events by region and by region-year), the system achieved 100% SQL accuracy across all 16 NERC subregions and both event types (64 queries total). These results validate the approach for target use cases, though performance on diverse natural language formulations requires further investigation. We discuss design trade-offs, failure modes including JSON output truncation, and pathways for extending this approach to other hazard domains.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]↗

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]↗

Agile Acceleration of LLVM Flang Support for Fortran 2018 Parallel Programming

The LLVM Flang compiler ("Flang") is currently Fortran 95 compliant, and the frontend can parse Fortran 2018. However, Flang does not have a comprehensive 2018 test suite and does not fully implement the static semantics of the 2018 standard. We are investigating whether agile software development techniques, such as pair programming and test-driven development (TDD), can help Flang to rapidly progress to Fortran 2018 compliance. Because of the paramount importance of parallelism in high-performance computing, we are focusing on Fortran’s parallel features, commonly denoted “Coarray Fortran.” We are developing what we believe are the first exhaustive, open-source tests for the static semantics of Fortran 2018 parallel features, and contributing them to the LLVM project. A related effort involves writing runtime tests for parallel 2018 features and supporting those tests by developing a new parallel runtime library: the CoArray Fortran Framework of Efficient Interfaces to Network Environments (Caffeine).

Rasmussen, Katherine↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗