Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Link Parser Library

Link Parser Library is a Python library for parsing HTTP Link header format. Link Parser extracts information about relationships to other resources from Link Header and translates it into an easy to use Json structure.

Balakireva, Lyudmila↗

Reposcanner

SAND2023-05455O Reposcanner provides a highly modular, extensible framework for defining routines for mining data from software repositories and performing analyses on that data to yield valuable insights on team behaviors. Reposcanner features seamless support for different version control platforms like GitHub, Gitlab, and Bitbucket; smart parsing of URLs; intelligent credential management capabilities; and a comprehensive test suite. Reposcanner is connected to the Exascale Computing Project and is intended for research purposes.

Mundt, Miranda↗

EMF Biomass Data Interpreter [SWR-23-06]

EMF Biomass Data Interpreter was developed for the transportation team of the Energy Modeling Forum (EMF) 37: Deep Decarbonization and High Electrification Scenarios for North America project. It parses results from different modeling groups to show how biomass might be employed under electrification.

Wachs, Elizabeth↗

Parsnip Parser Creation Application

Parsnip has three parts. The first part is the front-end user experience. The front end will be a graphical representation of the intermediate language. Once a user is satisfied with the information on the front end, Parsnip translates the data from the visual application into the second part of Parsnip - the intermediate language. More advanced users may skip the front end and generate their own intermediate language files. The final part is the backend which takes the intermediate language files and generates Zeek and Spicy code. Parsnip will not completely replace parser developers. Many protocols have unique challenges requiring manual effort; however, the goal of Parsnip is to automate at least 90% of the development that largely consists of repetitive tasks. Parsnip output will compile a functioning parser but may not include all PDU types or parse all data.

Huddleston, TimothyA.↗

CIMantic Graphs

CIMantic Graphs (aka CIM-Graph) is a new python library developed by PNNL to reduce the burden of working with the Common Information Model. CIMantic Graphs takes a novel approach of building in-memory labeled property graphs for creating, parsing, and editing CIM power system models.

Anderson, Alexander↗

RadSim: ENSDF, Xray and N42

This package includes three parts, (1) gov.bnl.nndc.ensdf, (2) gov.nist.xray and (3) gov.nist.physics.n42, which are utilized in the development of Radiation Detector Simulator (RadSim) project. RadSim is being developed to provide the capability to: (1) simulate radiation source emissions, (2) interpolate results from radiation transport tools into a common format to prepare incident flux, and (3) model radiation detector response to rapidly produce synthetic radiation measurement templates. RadSim is targeted for open-source release, which will enable researchers and industry partners to model gamma-ray detectors response to simulated flux from the transport tool of their choice. The first tool of the package, gov.bnl.nndc.ensdf, includes functionality to parse and split publicly available decay records in ENSDF format, and pull the relevant information from ENSDF libraries. The second tool, gov.nist.xray, provides Xray information using the public NIST database. Lastly, gov.nist.physics.n42, is a tool used to convert an N42 xml file into a Java object that can be interfaced within a Java program.

Hangal, DnaushA↗

FlowBench Raw Data Archive

The repo that provides data archive for DOE PoSeiDon project. It also contains scripts and instructions to parse the data.

George, Papadimitriou↗

Wind Energy Intrusion Detection System

SAND2022-14659 O The Wind Energy Intrusion Detection System uses a human-machine interface (HMI) to parse results and analyze for anomalies. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Johnson, Jay↗

Replete v. 2.0

SAND2021-15136 O Replete is a Java library that contains several common and useful utilities in a variety of categories including expression parsing, multi-threading, data pipelines, user interface, database connectivity, Unix shell emulation, information generalization, and Maven standardization. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

McClain, Jonathan↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

lanl-ansi/MG-RAVENS

The MG-RAVENS project with the DOE Office of Electricity Microgrid R&D Program is a project to develop a completely free, open-source data exchange standard (API) for the Department of Energy, targeted at software tools related to infrastructure modeling, particularly the modeling of microgrids and electric power distribution systems that are created with funding from the Microgrid R&D Program. This software produces formal definitions of an API, documentation, contains supporting functions for parsing, validating, etc., and will contain examples of workflows enabled by the developed API.

Fobes, David M↗

Library-AI-Toolset

Collection of tools designed to parse documents, such as PDFs, and extract structured elements including URLs, citation contexts, tables, formulas, and figures. This toolset leverages AI-based text extraction and classification methods, providing robust solutions for various scholarly resources processing needs.

Balakireva, Lyudmila↗

EyeON

EyeON: Eye on Operational technology Software Supply Chain attacks have risen drastically over the past few years, none more well-known and impactful than the SolarWinds compromise. Criminal organizations inserted an attack vector into a specific version of the source code, giving themselves an air of credibility. Once news broke on SolarWinds, identifying compromised sites was very difficult, even knowing the culprit update. Software Bills of Materials (SBOM) have been touted as the solution to reclaiming control of your software supply chain. Deployment of SBOMs has been slow, however, due to conflicting standards, opaque storage requirements, and vendor adoption. Additionally, the path from obtaining an SBOM and securing your supply chain is unclear; how can an SBOM library be leveraged to provide insight to your attack surface? The EyeON tool, sponsored by Department of Energy Cybersecurity, Energy Security, and Emergency Response (DoE CESER), aims to address these gaps by providing an encapsulated solution to tracking which updates have been installed in an enterprise, and alerting system administrators to vulnerabilities as they become known. Similar to a virus scanner, EyeON is a command line tool to parse either a single file, nested directory structure, or filesystem. It collects data such as signature (hashes), version information, VirusTotal tags, compiler, compilation date, and code signing information. Users will anonymously submit scan data periodically to DoE CESER, who will then compile a database of known software products employed by Critical Infrastructure and broadcast alerts based on discovered flaws as they arise.

Tenzing, Wangmo↗

test-grid-buildouts [SWR-23-112]

Test-grid-buildouts is a set of notebooks for creating buildouts of synthetic transmission grids with various levels of renewable energy. Workflows can be used to augment various test grid with varying levels of wind/solar generation. Aggregated renewable energy time series (actuals and scenarios) at grid buses can then be used for various transmission grid operation problems. Highly customizable workflow consists of parsing of synthetic grid files, setting up SAM & reV simulations, application of exclusion masks, creating of custom renewable generators on the grid, aggregating of power time series, and writing of modified grid description files.

Satkauskas, Ignas↗

pyscan-tlk

SAND2024-13867O pyscan-tlk software provides straight-forward access to control Thorlabs brand instruments with python. C bindings and python wrappers are generated and automatically based on text parsing the C documentation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mounce, Andrew↗

viral_fam_classifier

This code applies publicly available Hidden Markov Model (HMM) profiles to publicly available reference genomes of viruses. It parses the output of these searches to determine bitscore cut-offs for viral taxonomic lineages that may be represented by each HMM. Subsequently, the code will allow the user to classify novel viral sequences using these bitscore cut-offs. This code is a work-in-progress.

Kantor, Rose [Lawrence Livermore National Laborato↗

Ocpp 2.0.1. Interim Kpi Calculator

The project is split into four pieces. The first is a raw OCPP log parser. The second is a file splitter. The third is a message parser. The final piece is the Interim KPI calculator. The OCPP log parser was created from two different formats of raw OCPP 2.0.1 data. Its intended purpose is to extract device IDs and OCPP event messages from nontabular text logs. The parser looks for specific substrings in the logs to identify which of the two "standards" it should select from. The KPI generator does not perform any of its calculations in parallel. Instead, we opt for a naive batching approach. The splitter takes the file generated from the parser and creates many smaller files for each of the device IDs in the dataset. This allows the pandas queries in the log formatter to be iterate over a significantly smaller slice of data, increasing performance significantly. The message parser step takes messages from each of the files (containing distinct device IDs) and breaks the message out into pieces. The final result is a file with different columns specifying different attributes of the JSON message. The file is an aggregation of all different devices. This is the most complex portion of the code. The KPI calculator takes the parsed messages, as a single file, and calculates the KPI from that data. An excel file is produced with four sheets. These contain the metrics for Session Success, Charge Start Success, Charge End Success, and Charge Start Time. It includes the metrics for the different equations in the Interim KPI Implementation Guide as well as a weighted sum of the different equations for each KPI (excluding Charge End Success and Charge Start Time).

Quinn, Casey↗

GenConfig

SAND2025-04091O GenConfig converts a build name into a set of configuration flags or CMake fragment files for use with CMake. This is accomplished using ConfigKeywordParser and two configuration files. GenConfig is the main tool in a set of software libraries used for generating and configuring an environment and configuration flags. The tool uses other modules within the GenConfig family to ultimately parse and enable an environment that is ready for development from a given build name string. The unique algorithms used in GenConfig mainly pertain to validating the format and checking the existence of the given build string in the expected configuration .ini files. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Gates, Jason↗