Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Information Automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automation of Vulnerability and Patch Management: Information Extraction, Association, and Optimization

Vulnerability and patch management is an integral part of a robust cybersecurity program, yet it grows increasingly complex due to the sheer amount of data that must be analyzed. Particularly in Operational Technology (OT) environments, analysis must be done manually because of the lack of automated solutions. Additionally, there are many steps in this process, from the initial discovery of the vulnerability to the implementation of its remediation, and each step in the process requires different data in order to be performed effectively. In this work, we provide approaches and strategies to assist operators in industrial or OT environments throughout the vulnerability management cycle. Security advisories provide key information about mitigation strategies, or actions that can be taken when a patch is unavailable or cannot be installed. Details of these strategies are not shared in public vulnerability databases and must be found manually. We approach this problem by designing a solution to automatically identify that information within vendor security advisories and retrieve it for operator use. We start with an approach that requires domain-specific knowledge of certain frequently-seen reference websites. Next, an approach that can work on an arbitrary website but relies on certain keywords. Finally, an approach that uses Natural Language Processing (NLP) methods and does not require specific knowledge or keywords. Each of these approaches is more general than its predecessor; we demonstrate high accuracy for all approaches Advisories also often contain details of affected products in non-standard or natural language formats. While this information can be easily understood when read by an operator, the non-standard format acts as a barrier to effective automation. We provide an approach for the first step in this process: identifying vendors in security advisories and mapping them to a standard framework for representing digital assets and software products. We evaluate five established string similarity algorithms, plus one of our own design that combines string similarity and information theory, on the task of mapping vendors to their corresponding entries in the Common Platform Enumeration (CPE) repository. Our results show that our proposed metric outperforms all others. Due to the constraints on time, finances, and personnel for organizations, Large Language Models (LLMs) may seem like attractive opportunities for security operators to speed up information gathering; however, it is still not clear whether LLMs can handle vulnerability management tasks well. To answer this question, we perform an empirical study of LLMs’ ability to provide consistent, accurate information about vulnerabilities in order to guide organizations in their adoption of LLMs. We observe poor performance for all models tested, suggesting that these models are not well-suited to the consistent retrieval of accurate vulnerability information. Finally, once vulnerabilities have been identified and any additional information has been obtained, operators must decide which remediation actions to implement based on their available resources. This already-complex problem becomes even more so when we consider that a vulnerability may have multiple avenues for remediation. We formulate this scenario as two knapsack problems and provide solutions, which we then compare against several existing strategies for vulnerability prioritization seen in real operational environments.

McClanahan, Kylie↗

Exploiting correlations in multi-coincidence Coulomb explosion patterns for differentiating molecular structures using machine learning

Coulomb explosion imaging (CEI) is a powerful technique for capturing the real-time motion of individual atoms during ultrafast photochemical reactions. CEI generates high-dimensional data with naturally embedded correlations that allow mapping the coordinated motion of nuclei in molecules. This enables reliable separation of competing reaction pathways and makes this approach uniquely suited for characterizing weak reaction channels. However, rich information contained in experimental CEI patterns remains largely underexploited due to challenges in visualizing correlations between multiple observables in multi-dimensional parameter space. Here we present a new approach to CEI of intermediate-sized polyatomic molecules, detecting up to eight ionic fragments in coincidence and leveraging machine-learning-based analysis to identify patterns and correlations in the resulting high-dimensional momentum-space data, enabling robust molecular structure identification and differentiation. Our approach provides high-dimensional background-free data encoding exceptionally rich structural information and establishes an automated, scalable framework for extracting insightful information from the data. As a demonstration, we apply this method to image and distinguish dichloroethylene isomers, showcasing its potential for broader applications in molecular imaging. Our results pave the way for channel-specific analysis of ultrafast structural dynamics in chemically relevant systems, particularly for disentangling mixed reaction pathways and detecting contributions from weak channels and minority species.

Chemical Physics (physics.chem-ph)↗

Testing SOAR tools in use

Investigations within Security Operation Centers (SOCs) are tedious as they rely on manual efforts to query diverse data sources, overlay related logs, correlate the data into information, and then document results in a ticketing system. Security Orchestration, Automation, and Response (SOAR) tools are a relatively new technology that promise, with appropriate configuration, to collect, filter, and display needed diverse information; automate many of the common tasks that unnecessarily require SOC analysts’ time; facilitate SOC collaboration; and, in doing so, improve both efficiency and consistency of SOCs. There has been no prior research to test SOAR tools in practice; hence, understanding and evaluation of their effect is nascent and needed. Here, in this paper, we design and administer the first hands-on user study of SOAR tools, involving 24 participants and six commercial SOAR tools. Our contributions include the experimental design, itemizing six characteristics of SOAR tools, and a methodology for testing them. We describe configuration of a cyber range test environment, including network, user, and threat emulation; a full SOC tool suite; and creation of artifacts allowing multiple representative investigation scenarios to permit testing. We present the first research results on SOAR tools. Concisely, our findings are that: per-SOC SOAR configuration is extremely important; SOAR tools increase efficiency and reduce context switching, although with potentially decreased ticketing accuracy/completeness; user preference is slightly negatively correlated with their performance with the tool; internet dependence varies widely among SOAR tools; and balance of automation with assisting decision making is preferred by senior participants. We deliver a public user- and tool-anonymized and -obfuscated version of the data.

97 MATHEMATICS AND COMPUTING↗

Accessible Content Optimization for Research Needs (ACORN)

ACORN employs a set of automated processes for informing and/or enforcing defined content schemas to create standardized and highly structured data. Because of its standardized data source, ACORN easily applies computer automation to generate communication assets such as PDFs, Powerpoint presentations, and web pages. Built using the memory-safe Rust programming language, ACORN is portable and accessible for use on any Windows, Mac, or Linux machine.

Wohlgemuth, JasonHoward [Oak Ridge National Labora↗

Exploring the role of judgement and shared situation awareness when working with AI recommender systems

Abstract AI-advised Decision Making is a form of human-autonomy teaming in which an AI recommender system suggests a solution to a human operator, who is responsible for the final decision. This work seeks to examine the importance of judgement and shared situation awareness between humans and automated agents when interacting together in the form of a recommender systems. We propose manipulating both human judgement and shared situation awareness by providing the human decision maker with relevant information that the automated agent (AI), in the form of a recommender system, uses to generate possible courses of action. This paper presents the results of a two-phase between-subjects study in which participants and a recommender system jointly make a high-stakes decision. We varied the amount of relevant information the participant had, the assessment technique of the proposed solution, and the reliability of the recommender system. Findings indicate that this technique of supporting the human’s judgement and establishing a shared situation awareness is effective in (1) boosting the human decision maker’s situation awareness and task performance, (2) calibrating their trust in AI teammates, and (3) reducing overreliance on an AI partner. Additionally, participants were able to pinpoint the limitations and boundaries of the AI partner’s capabilities. They were able to discern situations where the AI’s recommendations could be trusted versus instances when they should not rely on the AI’s advice. This work proposes and validates a way to provide model-agnostic transparency into recommender systems that can support the human decision maker and lead to improved team performance.

Srivastava, Divya↗

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES↗

Utilization of traceable standards to validate plutonium isotopic purification and separation of plutonium progeny using AG MP-1M resin for nuclear forensic investigations

Radio-chronometric studies on plutonium (Pu) materials require independent measurement of the Pu (parent) content and isotopic distribution as well as concentration and isotopic distribution of the plutonium isotopic decay products. We performed a series of experiments to demonstrate the consistency of separations using the Lewatit MP 800 macroporous anion exchange resin and the AG MP-1M resin with traceable Pu isotopic certified reference material (CRM) standards 136, 137, 138, and 126-A. Two different mesh-sizes of the AG MP-1M resin were tested and the 50–100 mesh size resin was found to work more efficiently for the separation task. Both Lewatit and AG MP-1M resins were found to perform satisfactorily for quantitatively extracting the americium (Am) and uranium (U) progeny as well as gallium (Ga) present as a tracer in the Pu material. Both resins were effective in removing isobaric interferences from the Pu fraction used in isotopic measurements by thermal ionization mass spectrometry (TIMS). To address the co-elution of uranium and gallium, Alizarin red S (ARS) was used as a colorimetric dye to determine the behavior of UO 2 2+ and Ga 3+ on AG MP-1M resin with various acidic solutions as eluents using UV–vis spectra. Poor resolution of these peaks complicated quantitative analysis by UV–vis spectroscopy, but these results were informative in planning automated separation experiments by HPLC. LA-UR-24-28919.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Modeling Air Handling Units to Create a Diverse Fault Dataset for FDD Innovation: Lessons Learned and Recommendations

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault datasets for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling unit and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for the air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, two detailed AHU models, which included the single duct AHU and dual duct AHU developed in the Modelica language and HVACSIM+ were employed to carry out annual simulations of numerous common sensor faults, mechanical faults, and control sequence faults. The fault inclusive data were then validated by comparing fault effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. We report some lessons learnt during the efforts of validating the high volumes of the FDD data sets. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Development of a Annual Air Handling Unit Fault Dataset for FDD Tools: Lessons Learned and Considerations for FDD Developers

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault data for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling units and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for a single duct air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, a detailed AHU model was employed to carry out annual simulations of numerous common sensor and mechanical faults, which were then validated by comparing their effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

A rule-free workflow for the automated generation of databases from scientific literature

Abstract In recent times, transformer networks have achieved state-of-the-art performance in a wide range of natural language processing tasks. Here we present a workflow based on the fine-tuning of BERT models for different downstream tasks, which results in the automated extraction of structured information from unstructured natural language in scientific literature. Contrary to existing methods for the automated extraction of structured compound-property relations from similar sources, our workflow does not rely on the definition of intricate grammar rules. Hence, it can be adapted to a new task without requiring extensive implementation efforts and knowledge. We test our data-extraction workflow by automatically generating a database for Curie temperatures and one for band gaps. These are then compared with manually curated datasets and with those obtained with a state-of-the-art rule-based method. Furthermore, in order to showcase the practical utility of the automatically extracted data in a material-design workflow, we employ them to construct machine-learning models to predict Curie temperatures and band gaps. In general, we find that, although more noisy, automatically extracted datasets can grow fast in volume and that such volume partially compensates for the inaccuracy in downstream tasks.

36 MATERIALS SCIENCE↗

Field Validation of a Building Operating System Platform

The U.S General Services Administration's (GSA's) Green Proving Ground program, in partnership with the National Renewable Energy Laboratory. completed a large pilot study of an Energy Management Information Systems (EMIS) with Automated System Optimization (ASO). Four test bed facilities, each with different building characteristics and systems, were chosen for the implementation of cloud-based EMIS with ASO. Depending on functionality, this tool can be extremely effective in energy management and energy optimization in buildings. The capabilities evaluated in the pilot ranged from energy savings and energy consumption predictions to evaluations of user acceptance, operability, and ease of installation. This report presents the methodology, lessons learned and best practices, and deployment recommendations for the GSA's portfolio of commercial office space, comprising more than 8,500 properties.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Comparing Automated Posterior Estimation Techniques for Modeling Strong Lenses In Ground-based Survey Data

Current and future ground-based cosmological surveys, such as the Dark Energy Survey (DES), and the Vera Rubin Observatory Legacy Survey of Space and Time (LSST), are predicted to discover thousands to tens of thousands of strong gravitational lenses. The large number of strong lenses discoverable in future surveys will make strong lensing a highly competitive and complementary cosmic probe. However, conventional lens modeling techniques are unable to scale up to the sheer number of lenses that will be discovered through upcoming surveys. Therefore, the use of automated lens analysis techniques is necessary. We demonstrate that machine learning methods can be used to automate the inference of informative model posteriors of strong lensing systems in ground-based surveys with credible uncertainty estimation. We present two Simulation-Based Inference (SBI) approaches for lens parameter estimation of galaxy-galaxy lenses. We demonstrate applications of Neural Posteriors Estima tors (NPEs) and Bayesian Neural Network (BNNs) to automate the inference of a 12-parameter lensing system for DES-like ground-based imaging data. We apply a suite of diagnostics (e.g., posterior coverage and SBC) to validate the performance of our methods. We find that NPEs outperform the BNN, producing posterior distributions that are for the most part both more accurate and more precise; in particular, several source-light model parameters are systematically biased in the BNN implementation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A case study in contrastive learning information combination: Application to technical forensics of additive manufacturing filament source identification

Combination of information from disparate data sources into a single decision is a core challenge in many fields, including the field of technical forensics. Technical forensics (TF) utilizes technical characterization of questioned samples to determine properties of that sample; these properties are then used to infer information of forensic interest, such as provenance, age, or attribution. TF is utilized in traditional forensic applications, such as the attribution of material fragments from an explosive, and in nuclear forensic applications, such as the attribution of actinides which have been interdicted out of regulatory control. The challenge of combining information from disparate sources, described alternately by many terms including “Data Fusion” and “Data Integration”, is exacerbated in the technical forensics domain due to at least two factors: the challenge of interpreting each information source singularly, and the relatively small data set sizes available. Extensive literature exists attempting to combine technical forensics information sources, both in manual and automated processes. These attempts are often bespoke to the specific information sources (such as the bi-, tri-, or quad-isotope chart (Moody, Grant, and Hutcheon 2005)), with some emerging examples of simple early- and late- fusion (, respectively). Simultaneous to the information combination efforts described in the previous paragraph, the field of natural language processing attempted (and largely succeeded) in combining information from multiple non-technical information sources. The ecosystem of “multi-modal” language models, which can take text and images as input, and generate text and images as output, became large and diverse by 2025 (Khan et al. 2025). In a generalized sense, many of these methods are trained by learning neural networks which can convert raw text or images into a vector of numbers describing the text or image, hereafter called “embeddings” and the neural networks performing the conversion are called “embedders”. By using a separate embedder for text and images, finding coincident text and images (such as images with their captions), and optimizing the parameters of the embedders such that the embeddings for the text and the image are similar, the field has found a bridge between text and images (Girdhar et al. 2023). It is the contention of the authors of this report that this insight is not limited to text and images but instead can be extended to any modality which can be found coincidently. The subject of the rest of this report is the application of this method to example multi-modal technical forensic data. Some details about the data used in this report are not appropriate for this report, and are included in a companion report (PNNL-38669).

36 MATERIALS SCIENCE↗

Analyzing Risks of Virtual Private Network Connections

The use of Splunk for analyzing VPN logs is an effective approach for identifying vulnerabilities in network endpoints. Splunk, a powerful platform for searching, monitoring, and analyzing machine-generated data, enables organizations to aggregate VPN logs in real-time, providing insights into network activity, user behavior, and potential security risks. By indexing VPN traffic and authentication logs, security teams can track abnormal patterns such as multiple failed login attempts, unusual IP addresses, or unexpected changes in bandwidth usage, all of which could indicate potential vulnerabilities or breaches. With Splunk’s advanced search and reporting capabilities, users can create custom dashboards and alerts to detect suspicious activities. Automated searches can flag endpoints exhibiting unusual behavior, while correlation analysis can identify links between compromised devices and broader network vulnerabilities. In particular, Splunk's machine learning capabilities can be leveraged to predict and prevent threats by identifying trends that might otherwise be missed in traditional log analysis. This proactive approach to monitoring VPN logs allows for the early detection of security weaknesses, enabling rapid response and minimizing potential damage to network integrity. By enhancing endpoint visibility, Splunk plays a crucial role in securing remote connections and safeguarding sensitive information. Additionally, Splunk’s automation and alerting features allow teams to create custom workflows that notify them of vulnerable or misconfigured endpoints identified through Shodan. This synergy between Splunk’s log analysis and Shodan’s device intelligence enhances an organization’s ability to proactively identify and mitigate security risks, improving the overall resilience of their VPN infrastructure.

97 MATHEMATICS AND COMPUTING↗

Leveraging Radiofrequency Identification Success Beyond Hazardous Material Inventory Management at a National Laboratory

Effective inventory management can be overshadowed by conflicting priorities in organizational procedures, particularly in research-focused institutions such as national laboratories that handle expensive, delicate, and hazardous materials. Here, this study investigated the potential of radiofrequency identification (RFID) technology, currently used for hazardous chemical inventory, in applications with higher metal interference and absorption, specifically pressure release device (PRD) compliance and nuclear container management, at Lawrence Livermore National Laboratory (LLNL). This study was done to document best practices to enhance inventory identification speeds for inventory reconciliation and inventory recall and to explore optimal configurations for RFID implementation compared to traditional manual methods of equipment management. Tests were conducted to determine the ideal RFID tag orientation (read at angles of 0°, 90°, and 270°), various container layouts (linear, separated, curved, operational), and ID methods such as manual, barcode, and RFID performing three trials per method per orientation. Results indicated that 0° was the optimal read angle for minimizing metallic interference, and the operational and curved arrangements significantly outperformed the linear and separated configurations in read speed. 3D printed mounts were developed and tested, increasing the read range of the RFID reader by up to 235% in cases of high metallic interference. The RFID technology demonstrated an average speed increase of 65% over a simplified manual identification, which supports the conclusion that RFID is a more efficient method for large hazardous inventory management and equipment reconciliation. Additionally, capturing meta-data, such as location and date, can be used to query for inventory recall and automated updating of record information.

42 ENGINEERING↗

Curb Allocation and Pick-Up Drop-Off Aggregation for a Shared Autonomous Vehicle Fleet

Advances in information technologies and vehicle automation have birthed new transportation services, including shared autonomous vehicles (SAVs). Shared autonomous vehicles are on-demand self-driving taxis, with flexible routes and schedules, able to replace personal vehicles for many trips in the near future. The siting and density of pick-up and drop-off (PUDO) points for SAVs, much like bus stops, can be key in planning SAV fleet operations, since PUDOs impact SAV demand, route choices, passenger wait times, and network congestion. Unlike traditional human-driven taxis and ride-hailing vehicles like Lyft and Uber, SAVs are unlikely to engage in quasi-legal procedures, like double parking or fire hydrant pick-ups. In congested settings, like central business districts (CBD) or airport curbs, SAVs and others will not be allowed to pick up and drop off passengers wherever they like. This paper uses an agent-based simulation to model the impact of different PUDO locations and densities in the Austin, Texas CBD, where land values are highest and curb spaces are coveted. In this paper 18 scenarios were tested, varying PUDO density, fleet size and fare price. The results show that for a given fare price and fleet size, PUDO spacing (e.g., one block vs. three blocks) has significant impact on ridership, vehicle-miles travelled, vehicle occupancy, and revenue. A good fleet size to serve the region’s 80 core square miles is 4000 SAVs, charging a $1 fare per mile of travel distance, and with PUDOs spaced three blocks of distance apart from each other in the CBD.

Hunter, Christian B.↗