Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

157 records · Page 9

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Virtual Engineering Software Framework for Integrated Biomass Conversion Modeling

This presentation covers the design and implementation of a software tool to systematically connect computational models of unit operations to simulate an integrated process of low-temperature conversion of biomass to fuel. This virtual engineering (VE) software was designed with the overarching goal of connecting unit models written in various programming languages and requiring different computational resources within a single, flexible framework. The models and features currently considered for the VE library include mechanistic models for pretreatment, enzymatic hydrolysis, and aerobic bioreaction; high-fidelity computational fluid dynamics (CFD) simulations for enzymatic hydrolysis and aerobic bioreaction; and the capability to perform techno-economic analyses (TEA) using Aspen Plus, a commercial software package. The CFD models require access to high-performance computing (HPC) resources, so in addition to handling multiple programming languages and interfaces, the VE software must also be capable of interacting with an HPC scheduler to submit, run, and post-process jobs. Using the Python programming language, a new VE software package has been developed that contains functionality to manage the input-output communication between various unit models, schedule simulations to run on NREL's HPC and analyze those results, and interface with existing TEA software workflows. A Jupyter-notebook GUI was also created to solicit user input and provide documentation. In cases where multiple models for a particular unit-operation exist, selection between models is accomplished through a simple checkbox, with the appropriate inputs and outputs being parsed and converted seamlessly in the background. Each operation makes use of a different programming language, but the flow of information from pretreatment to enzymatic hydrolysis to bioreaction is managed with an intuitive, centralized file-communication strategy. In this talk, the programming approach and implementation details of the notebook are presented for multiple possibilities of the conversion process, including a demonstration of the ability to manage HPC resources. Additionally, an example of a sensitivity study of treatment parameters governing the overall conversion outcome is shown which highlights the ease of defining new problems using the VE Notebook workflow and leads into a discussion of ongoing work to enable outer-loop optimization studies.

biofuel↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Can Large Language Models Understand Intermediate Representations?

Intermediate Representations (IRs) are essential in compiler design and program analysis, yet their comprehension by Large Language Models (LLMs) remains underexplored. This paper presents a pioneering empirical study to investigate the capabilities of LLMs, including GPT-4, GPT-3, Gemma 2, LLaMA 3.1, and Code Llama, in understanding IRs. We analyze their performance across four tasks: Control Flow Graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Our results indicate that while LLMs demonstrate competence in parsing IR syntax and recognizing high-level structures, they struggle with control flow reasoning, execution semantics, and loop handling. Specifically, they often misinterpret branching instructions, omit critical IR operations, and rely on heuristic-based reasoning, leading to errors in CFG reconstruction, IR decompilation, and execution reasoning. The study underscores the necessity for IR-specific enhancements in LLMs, recommending fine-tuning on structured IR datasets and integration of explicit control flow models to augment their comprehension and handling of IR-related tasks.

Jiang, Hailong↗

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING↗

Characterization of Pinhole Collimators for High-Resolution Gamma Imaging of Irradiated Fuel

Post-irradiation examination (PIE) of nuclear fuels requires imaging tools capable of resolving isotopic and spatial features with high throughput. This project contributes to a proof-of-concept effort aimed at advancing gamma emission tomography (GET) by evaluating novel fine-aperture pinhole collimators. Two Rose’s metal collimators, 100 µm (20° acceptance angle) and 350 µm (30° acceptance angle), were prototyped and characterized for their effectiveness in transporting gamma rays through the pinhole aperture. To support data collection, a Python-based data acquisition system was developed to coordinate a rotation stage, linear stage, and CZT detector, reducing latency in high-rate gamma event logging to one second per acquisition. Queue-based file writing enabled seamless real-time data capture for count rates up to 35,000 counts per second (cps). List-mode parsing algorithms were implemented to differentiate single and simultaneous gamma interactions for future tomographic reconstruction. Detector response was evaluated in both spectroscopy and list mode acquisition methods across varying source distances to confirm absolute and collimator efficiencies. Preliminary efficiency figures suggest effective collimation of gamma-rays with energies below 700 keV, with ~4% residual intensity through the aperture for Cs-137. The impact of collimator geometry on image quality is currently being evaluated. This groundwork supports the ongoing development of a sub mm resolution cone-beam CT system for imaging fuel phantoms, an essential step toward improving GET efficiency and accelerating nuclear fuel qualification efforts.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Web-based Preprocessing and Visualization of 3D FIB Tomography Data for Nuclear Fuel Characterization

Three-dimensional (3D) focused ion beam (FIB) tomography enables reconstruction of internal nuclear fuel features that can't be fully evaluated through surface imaging alone. This capability supports characterization of fuel constituents and defects under thermal and irradiation conditions relevant to microreactor development. However, large tomography datasets can create data-handling, loading, and visualization challenges, especially when image-stack preparation and file conversion must be completed with separate tools. The Computational Ultraspatial Tomography Toolkit for High-Resolution Object Analysis Tools (CUTTRHOAT) is an open-source web application being developed to display FIB tomography datasets available through the Nuclear Research Data System (NRDS). The current alpha version requires prepared HDF5 datasets and has limited integrated data-preparation capabilities. This project improves CUTTHROAT by adding dataset-folder selection, automatic input detection, dataset scanning, missing-slice identification, blank-slice insertion, and image-stack-to-HDF5 conversion. Two applications will be compared: the baseline CUTTHROAT alpha workflow and the updated application containing the integrated data-handling and preprocessing functions. Evaluation will consider dataset detection accuracy, conversion success, loading time, rendering responsiveness, application stability, and user interaction. Preliminary results demonstrate successful loading of existing HDF5 files and converted image stacks, while testing also identified performance reductions caused by excessive blank-slice generation. The updated workflow reduces reliance on external preparation tools and supports more direct movement from image stacks to color-code 3D visualization. Future work includes refining missing-slice handling, integrating additional preprocessing functions, like a denoising feature, parsing TIFF metadata for automatic voxel scaling, and adding manual X, Y, and Z voxel-spacing inputs for PNG and JPEG.

36 - MATERIALS SCIENCE↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Database-Agnostic Log Analysis and Monitoring Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, US, IL; Fermilab]↗

The Building Adapter: Automatic Mapping of Commercial Buildings for Scalable Building Analytics

This project creates new solutions for the manual metadata mapping problem: the costly process of creating a match between a building’s sensor data streams and the inputs of a building analytics engine. This goal is achieved by creating and improving techniques for metadata inference: automatically constructing new contextual information for sensing and control points based on the sensor point names and the raw time series values. The objective is to enable vendors to apply building analytics to 90% of buildings with no manual mapping, and to 10% of buildings with a 90% reduction in manual mapping. These targets are set for all types of metadata required by current analytics engines, including type, location, equipment type, and other relationships. The outcome of this project is a suite of solutions to the manual mapping problem collectively called the Building Adapter that allows vendors to apply analytics engines to new buildings at a significantly reduced cost.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Significance of Aggregation Methods in Functional Group Modeling

The growth of forests and the feedbacks between forests and environmental changes are central issues in the planetary carbon cycle, global climate change, and basic plant ecology. A challenge to understanding both growth and feedbacks from local to global scales is that many critical metabolic processes vary among species. An innovation in solving this challenge is the recognition that species can be lumped into “functional groups” based on metabolic similarity, and these functional groups can then be studied in computational models that simulate ecosystem function. Despite the vast resources devoted to functional group studies and the progress made by them, an important logical and biological question has not been formally addressed, “How do the groupings alter the results of modeling studies?” To what extent do modeling results depend on the choices made in aggregating taxa into functional groups. Here, we consider the effects of using different aggregation strategies in simulating the carbon dynamics of a deciduous forest. Understanding the impacts that aggregation strategy has on efforts to simulate regional-to-global-scale forest dynamics offers insights into both ecosystem regulation and model function and addresses this central problem in the study of carbon dynamics.

54 ENVIRONMENTAL SCIENCES↗

Large Language Models (LLMs) for Energy Systems Research

The integration of Large Language Models (LLMs) in energy systems research promises transformative results, as demonstrated in this work, particularly in the realms of information retrieval and legal document analysis. We have developed a chat-based interface, specifically designed to query an extensive corpus of technical reports from the National Renewable Energy Laboratory (NREL). This interface capitalizes on the natural language processing capabilities of LLMs, providing future consumers of NREL research with a user-friendly platform to access and extract valuable information from technical documents, thus enhancing the dissemination of research to the public. In addition to information retrieval, we have employed LLMs to extract renewable energy siting ordinances from a variety of legal documents, a task traditionally driven by significant human labor. This automated extraction not only supports the ongoing development of the high-impact NREL siting ordinance database but also ensures the database's accuracy and comprehensiveness. Crucially, we have augmented the performance of LLMs through the integration of a decision tree framework, resulting in a substantial improvement in extraction accuracy. Comparative analysis with manual efforts has shown that this approach not only rivals but also significantly surpasses human accuracy, heralding increased reliability in legal document analysis for energy systems research. To democratize access to these advancements and foster collaborative research, we introduce the "Energy Language Model" (ELM), an open-source software package. ELM encapsulates the methodologies and tools developed in this work, providing researchers and practitioners with a robust toolkit to conduct similar analyses within their respective domains. Through these contributions, this work underscores the immense potential of LLMs in revolutionizing energy systems research, improving accuracy, efficiency, and accessibility in the field.

automation↗