Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Star–Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

Abstract We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam Subaru Strategic Program data. In our experiments, we evaluate the performance of a principal component analysis embedded GP predictive model against other machine-learning algorithms, including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

Resolved Dwarf Galaxy Searches within ~5 Mpc with the Vera Rubin Observatory and Subaru Hyper Suprime-Cam*

We present a preview of the faint dwarf galaxy discoveries that will be possible with the Vera C. Rubin Observatory and Subaru Hyper Suprime-Cam in the next decade. In this work, we combine deep ground-based images from the Panoramic Imaging Survey of Centaurus and Sculptor (PISCeS) and extensive image simulations to investigate the recovery of faint, resolved dwarf galaxies in the Local Volume with a matched-filter technique. We adopt three fiducial distances - 1.5, 3.5, 5 Mpc, and quantitatively evaluate the effects on dwarf detection of varied stellar backgrounds, ellipticity, and Milky Way foreground contamination and extinction. We show that our matched-filter method is powerful for identifying both compact and extended systems, and near-future surveys will be able to probe at least ~4.5 mag below the tip of the red giant branch (TRGB) for a distance of up to 1.5 Mpc, and ~2 mag below the TRGB at 5 Mpc. This will push the discovery frontier for resolved dwarf galaxies to fainter magnitudes, lower surface brightnesses, and larger distances. Our simulations show the secure census of dwarf galaxies down to $M_{V}$$\approx$-5, -7, -8, will be soon within reach, out to 1.5 Mpc, 3.5 Mpc, and 5 Mpc, respectively, allowing us to quantify the statistical fluctuations in satellite abundances around hosts, and parse environmental effects as a function of host properties.

79 ASTRONOMY AND ASTROPHYSICS↗

Phase Doppler Interferometry for Efficient Cloud Drop Size Distribution, Number Density, and LWC Measurements

Threats to aviation safety as a result of super-cooled large drops (SLD) has been addressed by the FAA rules change (14 CFR Part 25) with the additional icing certification requirement. SLD clouds often consist of bi-modal drop size spectra leading to significant problems in simulating and characterizing these conditions in situ and in icing wind tunnels. Legacy instrumentation for measuring drop size distributions and liquid water content are challenged under these conditions. The large size range measurement problem is addressed with the development of the Phase Doppler Interferometer, Flight Probe Dual-Range (PDI FPDR). The method is described in this report along with the measurement capabilities including the dynamic measurement range and overall working size range. The PDI instrument bases drop size measurements on the light wavelength as the measurement length scale. The light wavelength is a much more robust scale, especially as compared to the light scattering intensity. Additionally, methods for accurately characterizing the sample volume in situ based on measured drop velocity and transit time are reviewed, given the importance of this parameter for merging results and measuring LWC. Droplet coincidence in the sample volume can be problematic so this condition is treated with an innovative signal parsing approach. Measurement examples acquired in the NASA IRT are provided. Measurements of LWC showed good agreement with the Artium Particle Imaging (PI) instrument but diverged from the tunnel calibration results for larger MVD values.

42 ENGINEERING↗

A novel methodology for assessing the hygroscopicity of aerosol filter samples

Abstract. Due to US regulations, concentrations of hygroscopic inorganic sulfate and nitrate have declined in recent years, leading to an increased importance of the hygroscopic nature of organic matter (OM). The hygroscopicity of OM is poorly characterized because only a fraction of the multitude of organic compounds in the atmosphere is readily measured, and there is limited information on their hygroscopic behaviors. Hygroscopicity of aerosol is traditionally measured using a humidified tandem differential mobility analyzer (HTDMA) or electrodynamic balance (EDB). EDB measures water uptake by a single particle. For ambient and chamber studies, HTDMA measurements provide water uptake and particle size information but not chemical composition. To fill this information gap, we developed a novel methodology to assess the water uptake by particles collected on Teflon filters. This method uses the same filter sample for both hygroscopicity measurements and chemical characterization, thereby providing an opportunity to link the measured hygroscopicity with ambient particle composition. To test the method, hygroscopic measurements were conducted in the laboratory for ammonium sulfate, sodium chloride, glucose, and malonic acid, which were collected on 25 mm Teflon filters using an aerosol generator and sampler. Constant-humidity solutions (CHSs), including potassium chloride, barium chloride dihydrate, and potassium sulfate, were employed in a saturated form to maintain the relative humidity (RH) at approximately 84 %, 90 %, and 97 % in small chambers. Our preliminary experiments revealed that, without the pouch, water uptake measurements were not feasible due to rapid water loss during weighing. Additionally, we observed some absorption by the aluminum pouch itself. To account for this, concurrent measurements were conducted for both the loaded and the blank filters at each RH level. Thus, the dry loaded and blank Teflon filters were placed in aluminum pouches with one side open and in RH-controlled chambers for more than 24 h. The wet loaded samples and wet blanks were then weighed using an ultramicrobalance to determine the water uptake by the respective compound and the blank Teflon filter. The net amount of water absorbed by each compound was calculated by subtracting the water uptake of the blank filter from that of the wet loaded filter. Hygroscopic parameters, including the water-to-solute (W / S) ratio, molality, mass fraction solute (mfs), and growth factors (GFs), were calculated from the measurements. The results obtained are consistent with those reported by the Extended Aerosol Inorganics Model (E-AIM) and previous studies utilizing HTDMA and EDB for these compounds, highlighting the accuracy of this new methodology. This new approach enables the hygroscopicity and chemical composition of individual filter samples to be assessed so that in complex mixtures, such as chamber and ambient samples, the total water uptake can be parsed between the inorganic and organic components of the aerosol.

54 ENVIRONMENTAL SCIENCES↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Cleaned 5-Minute Resolution Air Quality and Meteorological Data from Nine TCEQ CAMS Sites in Houston, Texas (Nov 2021 – Oct 2022)

These data encompass 5-minute air monitoring and meteorological observations collected in the greater Houston, Texas metropolitan region, at nine (9) Continuous Ambient Monitoring Stations (CAMS) operated by the Texas Commission on Environmental Quality (TCEQ) between November 1, 2021 and October 31, 2022. The CAMS sites (CAMS 1, 8, 35, 45, 148, 403, 405, 410, and 1052) were chosen because their instrumentation includes measurements of PM2.5. These sites also provide continuous multi-parameter air-quality and meteorological measurements. Particulate matter (PM2.5, PM10) was sampled along with several trace gases, including ozone (O3), nitrogen oxides (NO, NO2, NOx), sulfur dioxide (SO2), and carbon monoxide (CO). The data set also contains standard surface meteorological parameters (temperature, humidity, pressure, wind speed, and wind direction). Several sites also include AutoGC-based measurements of volatile organic compounds (VOCs). Air monitoring instruments deployed at the selected sites comprise the following systems: BAM-1020 or TEOM (PM2.5), Thermo Scientific TEI 49i (O3), TEI 42i (NOx), and AutoGCs (VOCs). This data set is similar to the data included within the houairq5mX1.00 datastream, except for a few additional quality control steps. A systematic data cleaning and verification process was performed on the data set to ensure its quality and preparation for analysis. Removal of non-numeric status flags (e.g., [LIM], [QAS], [SPZ], [CAL], [PMA], [AQI], [SPN], [MAL]) was accomplished by employing rule-based string parsing to extract valid numerical values. Missing entries were set to -9999; however, invalid or anomalous values (e.g., 99999) were retained as originally reported by the TCEQ to preserve data provenance. The time sequence was verified for completeness, removal of duplicates, and uniformity at 5-minute intervals. Column labeling was standardized, and corresponding values were assessed for physical plausibility. All timestamps in the data set were reported in Coordinated Universal Time (UTC) as provided by the TCEQ. Further, the latitude and longitude coordinates were added for each CAMS site. A subset of the data (June 1–September 30, 2022) has been used in the following publication: Subba et al. 2025. “Implications of sea breeze circulations on boundary layer aerosols in the southern coastal Texas region.” EGUsphere 2025: 1–49, https://doi.org/10.5194/egusphere-2025-2659.

latitude↗

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs↗

Virtual Engineering Software Framework for Integrated Biomass Conversion Modeling

This presentation covers the design and implementation of a software tool to systematically connect computational models of unit operations to simulate an integrated process of low-temperature conversion of biomass to fuel. This virtual engineering (VE) software was designed with the overarching goal of connecting unit models written in various programming languages and requiring different computational resources within a single, flexible framework. The models and features currently considered for the VE library include mechanistic models for pretreatment, enzymatic hydrolysis, and aerobic bioreaction; high-fidelity computational fluid dynamics (CFD) simulations for enzymatic hydrolysis and aerobic bioreaction; and the capability to perform techno-economic analyses (TEA) using Aspen Plus, a commercial software package. The CFD models require access to high-performance computing (HPC) resources, so in addition to handling multiple programming languages and interfaces, the VE software must also be capable of interacting with an HPC scheduler to submit, run, and post-process jobs. Using the Python programming language, a new VE software package has been developed that contains functionality to manage the input-output communication between various unit models, schedule simulations to run on NREL's HPC and analyze those results, and interface with existing TEA software workflows. A Jupyter-notebook GUI was also created to solicit user input and provide documentation. In cases where multiple models for a particular unit-operation exist, selection between models is accomplished through a simple checkbox, with the appropriate inputs and outputs being parsed and converted seamlessly in the background. Each operation makes use of a different programming language, but the flow of information from pretreatment to enzymatic hydrolysis to bioreaction is managed with an intuitive, centralized file-communication strategy. In this talk, the programming approach and implementation details of the notebook are presented for multiple possibilities of the conversion process, including a demonstration of the ability to manage HPC resources. Additionally, an example of a sensitivity study of treatment parameters governing the overall conversion outcome is shown which highlights the ease of defining new problems using the VE Notebook workflow and leads into a discussion of ongoing work to enable outer-loop optimization studies.

biofuel↗

Improving Cyber Situational Understanding

Effective cybersecurity operations require the ability to analyze large amounts of information to assess security risks and formulate defensive strategies against adversaries. This has become more complex in recent years as the sprawl and interconnectivity of devices grows through implementation of virtualization, cloud computing, and Internet of Things (IoT). The amount of data and analysis required for effective cybersecurity command and control decisions far exceeds humans’ capacity to perform manually. We characterize the analysis problem as cyber situational understanding. The research presented to improve cyber situational understanding focuses on vulnerability analysis and threat intelligence. Regarding vulnerabilities, entities must analyze and plan work for between thousands and tens of thousands of software vulnerabilities annually. Entities heavily use network firewalls to limit vulnerability exposure. As a result, some of these vulnerabilities permit exposure to adversarial exploitation, whereas others are inaccessible and therefore present negligible risk of exploitation. Distinguishing between high and low risk software vulnerabilities requires a deep understanding of the vulnerability, network firewall protection, and characteristics of the targeted device. This problem is solved by extracting network service features from vulnerability data features using both machine-learning and natural language processing. Then, the network firewall topology is parsed to determine which vulnerabilities are reachable by adversaries. Ultimately, a state-based safety analysis ascertains which vulnerabilities are unsafe. A related vulnerability analysis problem occurs in cybersecurity operations when associating an entity’s hardware and software assets to public vulnerability databases. Assets often reveal hardware and software through installation artifacts and network service identification, and entities store these artifacts in inventory databases. However, software and hardware vendors apply a standard Common Platform Enumeration (CPE) naming convention when publicly reporting vulnerabilities. Associating these two datasets often requires many hours to days of manual inspection. The proposed solution automates the mapping approach of human analysts using fuzzy matching techniques, natural language processing, and, ultimately, machine learning to present a small set of recommendations for mapping the two datasets. The result significantly reduces human analysis time and reduces the occurrence of false positives in vulnerability notifications. Finally, cyber threat intelligence (CTI) requires associating cyber observable artifacts, such as IP addresses, URIs, and file hashes, with cyber threat tactics, techniques, and procedures. Unfortunately, most CTI data is compartmentalized across multiple organizations and cannot be shared due to the legal and reputational risk with cyber threat being associated with the entity. The approach to solving this problem inovlves using a distributed ledger with anonymous token spending and authentication. This allows a consortium of semi-trusted entities to share the workload of curating CTI for a threat sharing community’s cooperative benefit.

Huff, Philip↗

Remote Instrumentation and Data Acquisition: An Internship Research Report

This report outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope’s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and lessons learned, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Fermilab]↗

Can Large Language Models Understand Intermediate Representations?

Intermediate Representations (IRs) are essential in compiler design and program analysis, yet their comprehension by Large Language Models (LLMs) remains underexplored. This paper presents a pioneering empirical study to investigate the capabilities of LLMs, including GPT-4, GPT-3, Gemma 2, LLaMA 3.1, and Code Llama, in understanding IRs. We analyze their performance across four tasks: Control Flow Graph (CFG) reconstruction, decompilation, code summarization, and execution reasoning. Our results indicate that while LLMs demonstrate competence in parsing IR syntax and recognizing high-level structures, they struggle with control flow reasoning, execution semantics, and loop handling. Specifically, they often misinterpret branching instructions, omit critical IR operations, and rely on heuristic-based reasoning, leading to errors in CFG reconstruction, IR decompilation, and execution reasoning. The study underscores the necessity for IR-specific enhancements in LLMs, recommending fine-tuning on structured IR datasets and integration of explicit control flow models to augment their comprehension and handling of IR-related tasks.

Jiang, Hailong↗

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING↗

Characterization of Pinhole Collimators for High-Resolution Gamma Imaging of Irradiated Fuel

Post-irradiation examination (PIE) of nuclear fuels requires imaging tools capable of resolving isotopic and spatial features with high throughput. This project contributes to a proof-of-concept effort aimed at advancing gamma emission tomography (GET) by evaluating novel fine-aperture pinhole collimators. Two Rose’s metal collimators, 100 µm (20° acceptance angle) and 350 µm (30° acceptance angle), were prototyped and characterized for their effectiveness in transporting gamma rays through the pinhole aperture. To support data collection, a Python-based data acquisition system was developed to coordinate a rotation stage, linear stage, and CZT detector, reducing latency in high-rate gamma event logging to one second per acquisition. Queue-based file writing enabled seamless real-time data capture for count rates up to 35,000 counts per second (cps). List-mode parsing algorithms were implemented to differentiate single and simultaneous gamma interactions for future tomographic reconstruction. Detector response was evaluated in both spectroscopy and list mode acquisition methods across varying source distances to confirm absolute and collimator efficiencies. Preliminary efficiency figures suggest effective collimation of gamma-rays with energies below 700 keV, with ~4% residual intensity through the aperture for Cs-137. The impact of collimator geometry on image quality is currently being evaluated. This groundwork supports the ongoing development of a sub mm resolution cone-beam CT system for imaging fuel phantoms, an essential step toward improving GET efficiency and accelerating nuclear fuel qualification efforts.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Web-based Preprocessing and Visualization of 3D FIB Tomography Data for Nuclear Fuel Characterization

Three-dimensional (3D) focused ion beam (FIB) tomography enables reconstruction of internal nuclear fuel features that can't be fully evaluated through surface imaging alone. This capability supports characterization of fuel constituents and defects under thermal and irradiation conditions relevant to microreactor development. However, large tomography datasets can create data-handling, loading, and visualization challenges, especially when image-stack preparation and file conversion must be completed with separate tools. The Computational Ultraspatial Tomography Toolkit for High-Resolution Object Analysis Tools (CUTTRHOAT) is an open-source web application being developed to display FIB tomography datasets available through the Nuclear Research Data System (NRDS). The current alpha version requires prepared HDF5 datasets and has limited integrated data-preparation capabilities. This project improves CUTTHROAT by adding dataset-folder selection, automatic input detection, dataset scanning, missing-slice identification, blank-slice insertion, and image-stack-to-HDF5 conversion. Two applications will be compared: the baseline CUTTHROAT alpha workflow and the updated application containing the integrated data-handling and preprocessing functions. Evaluation will consider dataset detection accuracy, conversion success, loading time, rendering responsiveness, application stability, and user interaction. Preliminary results demonstrate successful loading of existing HDF5 files and converted image stacks, while testing also identified performance reductions caused by excessive blank-slice generation. The updated workflow reduces reliance on external preparation tools and supports more direct movement from image stacks to color-code 3D visualization. Future work includes refining missing-slice handling, integrating additional preprocessing functions, like a denoising feature, parsing TIFF metadata for automatic voxel scaling, and adding manual X, Y, and Z voxel-spacing inputs for PNG and JPEG.

36 - MATERIALS SCIENCE↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Database-Agnostic Log Analysis and Monitoring Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, US, IL; Fermilab]↗

The Building Adapter: Automatic Mapping of Commercial Buildings for Scalable Building Analytics

This project creates new solutions for the manual metadata mapping problem: the costly process of creating a match between a building’s sensor data streams and the inputs of a building analytics engine. This goal is achieved by creating and improving techniques for metadata inference: automatically constructing new contextual information for sensing and control points based on the sensor point names and the raw time series values. The objective is to enable vendors to apply building analytics to 90% of buildings with no manual mapping, and to 10% of buildings with a 90% reduction in manual mapping. These targets are set for all types of metadata required by current analytics engines, including type, location, equipment type, and other relationships. The outcome of this project is a suite of solutions to the manual mapping problem collectively called the Building Adapter that allows vendors to apply analytics engines to new buildings at a significantly reduced cost.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗