Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

BL7011 v1.0.0

This is a software package tailored to the needs of beam line 7.0.1.1 at the Advanced Light Source. It use open source packages to increase workflow efficiency around the beam line. This software package provides analysis code to deal with tasks connected to coherent x-ray measurements at the Cosmic Scattering end station.

Morley, Sophie [Lawrence Berkeley National Laborat↗

SMART Mobility. Modeling Workflow Development, Implementation, and Results Capstone Report

The U.S. Department of Energy’s Systems and Modeling for Accelerated Research in Transportation (SMART) Mobility Consortium is a multiyear, multi-laboratory collaborative, managed by the Energy Efficient Mobility Systems Program of the Office of Energy Efficiency and Renewable Energy, Vehicle Technologies Office, dedicated to further understanding the energy implications and opportunities of advanced mobility technologies and services. The first three-year research phase of SMART Mobility occurred from 2017 through 2019, and included five research pillars: Connected and Automated Vehicles, Mobility Decision Science, Multi-Modal Freight, Urban Science, and Advanced Fueling Infrastructure. A sixth research thrust integrated aspects of all five pillars to develop a SMART Mobility Modeling Workflow to evaluate new transportation technologies and services at scale. This report summarizes the work of the SMART Mobility Modeling Workflow effort. The SMART Mobility Modeling Workflow was developed to evaluate new transportation technologies such as connectivity, automation, sharing, and electrification through multi-level systems analysis that captures the dynamic interactions between technologies. By integrating multiple models across different levels of fidelity and scale, the Workflow yields insights about the influence of new mobility and vehicle technologies at the system level. For information about the other Pillars, please refer to the relevant pillar’s Capstone Report.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs

It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective use of allocations, and to co-locate compute jobs with large datasets. Unfortunately, moving scientific applications between HPC systems is challenging for various reasons, most notably that HPC systems have different HPC schedulers. We introduce PSI/J, a job management abstraction API intended to simplify the construction of software components and applications that are portable over various HPC scheduler implementations. We argue that such a system is both necessary and that no viable alternative currently exists. We analyze similar notable APIs and attempt to determine the factors that influenced their evolution and adoption by the HPC community. We base the design of PSI/J on that analysis. We describe how PSI/J has been integrated in three workflow systems and one application, and also show via experiments that PSI/J imposes minimal overhead.

Hategan, Mihael↗

Use of Longitudinal Serum Analysis and Machine Learning to Develop a Classifier for Cancer Early Detection

Early detection of solid tumors through a simple screening process, such as the proteomic analysis of biofluids, has the potential to significantly alter the management and outcomes of cancers. The application of advanced targeted proteomics measurements and data analysis strategies to uniformly collected serum or plasma samples would enable longitudinal studies of cancer risk, progression, and response to therapy that have the potential to significantly reduce cancer burden in general. In this article, we describe a generalizable workflow combining robust, multiplexed targeted proteomics measurements applied to longitudinal samples from the Department of Defense Serum Repository with a Random Forest machine learning method for developing and initially evaluating the performance of candidate biomarker panels for early detection of cancers. The effectiveness of this approach was demonstrated in a cohort of 175 head and neck squamous cell carcinoma patients. The outlined protocols include methods for sample preparation, instrument analysis, and data analysis and interpretation using this workflow.

Longitudinal analysis, machine learning, cancer, e↗

Standardized phylogenetic and molecular evolutionary analysis applied to species across the microbial tree of life

There is growing interest in reconstructing phylogenies from the copious amounts of genome sequencing projects that target related viral, bacterial or eukaryotic organisms. To facilitate the construction of standardized and robust phylogenies for disparate types of projects, we have developed a complete bioinformatic workflow, with a web-based component to perform phylogenetic and molecular evolutionary (PhaME) analysis from sequencing reads, draft assemblies or completed genomes of closely related organisms. Furthermore, the ability to incorporate raw data, including some metagenomic samples containing a target organism (e.g. from clinical samples with suspected infectious agents), shows promise for the rapid phylogenetic characterization of organisms within complex samples without the need for prior assembly.

59 BASIC BIOLOGICAL SCIENCES↗

SBbadger: biochemical reaction networks with definable degree distributions

Abstract Motivation An essential step in developing computational tools for the inference, optimization and simulation of biochemical reaction networks is gauging tool performance against earlier efforts using an appropriate set of benchmarks. General strategies for the assembly of benchmark models include collection from the literature, creation via subnetwork extraction and de novo generation. However, with respect to biochemical reaction networks, these approaches and their associated tools are either poorly suited to generate models that reflect the wide range of properties found in natural biochemical networks or to do so in numbers that enable rigorous statistical analysis. Results In this work, we present SBbadger, a python-based software tool for the generation of synthetic biochemical reaction or metabolic networks with user-defined degree distributions, multiple available kinetic formalisms and a host of other definable properties. SBbadger thus enables the creation of benchmark model sets that reflect properties of biological systems and generate the kinetics and model structures typically targeted by computational analysis and inference software. Here, we detail the computational and algorithmic workflow of SBbadger, demonstrate its performance under various settings, provide sample outputs and compare it to currently available biochemical reaction network generation software. Availability and implementation SBbadger is implemented in Python and is freely available at https://github.com/sys-bio/SBbadger and via PyPI at https://pypi.org/project/SBbadger/. Documentation can be found at https://SBbadger.readthedocs.io. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Demonstration of Model-Based Design for Digital Controller Using Formal Methods

This report describes work originally performed in FY19 that assembled a workflow enabling formal verification of high-consequence digital controllers. The approach builds on an engineering analysis strategy using multiple abstraction levels (Model-Based Design) and performs exhaustive formal analysis of appropriate levels – here, state machines and C code – to assure always/never properties of digital logic that cannot be verified by testing alone. The operation of the workflow is illustrated using example models and code, including expected failures of verification when properties are violated.

97 MATHEMATICS AND COMPUTING↗

Machine learning for automated experimentation in scanning transmission electron microscopy

Abstract Machine learning (ML) has become critical for post-acquisition data analysis in (scanning) transmission electron microscopy, (S)TEM, imaging and spectroscopy. An emerging trend is the transition to real-time analysis and closed-loop microscope operation. The effective use of ML in electron microscopy now requires the development of strategies for microscopy-centric experiment workflow design and optimization. Here, we discuss the associated challenges with the transition to active ML, including sequential data analysis and out-of-distribution drift effects, the requirements for edge operation, local and cloud data storage, and theory in the loop operations. Specifically, we discuss the relative contributions of human scientists and ML agents in the ideation, orchestration, and execution of experimental workflows, as well as the need to develop universal hyper languages that can apply across multiple platforms. These considerations will collectively inform the operationalization of ML in next-generation experimentation.

36 MATERIALS SCIENCE↗

Land-use analysis using infrastructure representations and high-resolution flood inundation mapping techniques

In the face of climate change and population growth in coastal regions, land-use analysis efforts are more challenging than ever. Land-use decision-makers in coastal communities are burdened with the difficult choices of where to place new homes versus other assets. While there has been an increased focus on hazard mitigation and disaster resilience in the field of planning, evidence points towards continued development in risk-prone areas including flood zones. Residential development within flood zones specifically continues to be a major issue. To help counter this trend, this study introduces a novel land-use analysis method, coupling topographic flood inundation mapping techniques with digital elevation model (DEM) adaptations. This Topographic Model Scenario Generation workflow can be used by planners early in the land-use decision making process and provides an alternative to high-computational hydraulic models. The analysis also includes the identification of strengths and weaknesses of topographic models' recognition of built infrastructure assets, adding to a limited body of knowledge addressing recommended uses of such models. Levees and canals prove particularly functional in this context while detention ponds less so, likely due to a lack of total water mass accountability. Lastly, we provide a functional demonstration in Southeast Texas to illustrate the workflow's ability to create multiple infrastructure scenarios and visualize their effects across different flood events.

42 ENGINEERING↗

PipeSight: A High-Performance Computing Platform for Pipeline Integrity Management

The Phase I feasibility study completed as part of this project has led to a number of innovative technologies being developed and has laid the foundation for a successful Phase II effort to commercialize a platform for managing the integrity of pipelines for the damage mechanisms of the new, hybrid-energy based economy. To ground the development efforts and direction of the project, an extensive market research and customer discovery effort was undertaken early in Phase I. Through this effort, a number of pipeline owners and operators were interviewed, and the following key findings were discovered about the pipeline industry: • Small pipeline operators do not have the central engineering groups necessary to perform their own independent analysis of inspection data, but instead rely on summarized tally sheets provided to them by inspection service providers. • The time it takes to go from an inspection to a completed engineering assessment, even for small segments of pipeline, can take anywhere from 30-120 days. During this delay, critical threats can (and have been known to) cause failures. • Uncertainty is often not accounted for in the assessment of pipeline integrity. The tally sheets provided by third-party service providers are almost always deterministic in nature, identifying threats that present a concern only to the current (not the future) integrity of the pipeline. • It is uncommon to apply the latest technologies to perform advanced assessments of damaged pipelines. There is a desire to use more advanced analysis capabilities to assess threats. Many pipeline operators indicated that they would often excavate a pipeline to perform an inspection and find that the damage was not as bad as they anticipated, thus using limited resources unnecessarily. Companies are not consistent in their use of inspection data to determine corrosion rates, and those that do only calculate deterministic corrosion rates. • The industry has prominently relied on time-based inspections but has recently started to transition to risk-based inspections. However, there appears to be no uniform guidance on how to do so while properly accounting for all sources of uncertainty. • Companies are not storing inspection data in a manner that allows for the ready determination of temporal trends. • Predictive maintenance principles and practices are beginning to be used by early adopters • Some pipelines are being re-purposed to transport different process fluids than they were designed for, e.g., H 2 and CO 2 rich process streams to serve the new hybrid-energy based economy, which are presenting new integrity concerns for the existing pipeline network that crisscrosses the United States. As a result of these discoveries, we were able to target the development efforts in Phase I to best serve the needs of the industry. In Phase I, we developed a way to correlate multiple large-scale scans of the pipeline to determine a probabilistic corrosion rate that accounts for all sources of error and uncertainty in the inspection process. This probabilistic corrosion rate can be used to predict the future thickness distribution of the pipe wall. We demonstrate how this analysis may be performed in an analytical fashion and has been implemented in such a manner that it can be readily distributed using GPU computing through integration of the Kokkos programming model. We also make a very novel extension of the analytical corrosion rate model to Bayesian Networks (an explainable AI technique) that can account for non-parametric distributions of corrosion rates. With the predictions made above for the probabilistic corrosion rate and corresponding future distribution of the pipe wall thickness, we can assess the integrity of the pipeline through the use of a probabilistic engineering assessment. We developed a novel screening data analysis approach that can rapidly identify ‘hotspots’ (local thin areas) where the integrity of the pipeline is a concern. Once more, we implemented this screening approach in C++ to leverage GPU computing via the Kokkos programming model. After the critical hotspots are identified, we developed a program that can automatically generate an advanced finite element model of the damaged regions. Since the number of damaged regions that require advanced analysis can number in the thousands, we integrated an open-source container-native workflow engine for orchestrating parallel jobs on the cloud. Initially, these advanced numerical models were only designed to account for loading due to internal pressure. However, in a slight pivot from the initial Phase I proposal, we developed a complete pipe stress analysis program (called Simflex) which can simulate the complete pipeline and its response to thermal expansion, pressure, thermal bowing, weight, wind, earthquake, support displacement, support friction and external forces. This pipe stress analysis program was written generically, to handle any piping system, but contains the features needed to model long pipelines (i.e., it incorporates a model for soil mechanics and can account for the nonlinear boundary conditions necessary to simulate long underground pipelines). This pipe stress analysis program can simulate any segment of the pipeline (simple or complex) under any set of conditions and loads, to determine the supplemental loads (axial forces and bending moments) at the location of damage. This enables the most accurate state of stress to be accounted for in the pipeline, which can prove critical when evaluating the integrity of a damaged region. In the process of developing the technologies to perform the integrity assessment of the pipeline, we also extended one of the industry standard approaches for performing the assessment of local thin areas that extend more in the circumferential direction than the longitudinal direction of the pipeline. This approach was presented to the API 579-1/AS ME FFS-1 steering committee in November 2021 for consideration in the next edition of the industry standard for Fitness-For-Service (expected to be released in 2023). To help pipeline operators make decisions with the results on any integrity assessment, we developed a new approach to the life-cycle management of pipelines which uses a Bayesian Decision Network. The network is designed to help pipeline operators plan and prioritize inspection activities and ultimately make smarter, more cost-effective decisions. The Bayesian approach accounts for all sources of uncertainty and carries them through to the final optimal decisions, providing a probabilistic framework for optimizing inspection intervals. The proof-of-concept networks developed in the feasibility study are complete, verified, and are focused on a subset of the pipeline. To expand this novel approach to the scale necessary for an entire network of pipelines in Phase II, we will leverage the DOE-funded Bengi solver for industrial-scale decision making with Bayesian Networks [22]. Once implemented, we will be able to provide the pipeline industry with a much-needed tool for optimal inspection planning using truly explainable artificial intelligence (XAI). To handle all of these advanced capabilities into a cloud-based platform, the architecture of the Equity Engineering Cloud (EEC) was extended to include Argo Workflows, a framework capable of distributing and managing a massive number of jobs that consume their own resources, such that thousands of serial finite element simulations can be run in parallel. As part of this substantial undertaking, we also integrated Argo Continuous Delivery (CD) into the EEC, to aid with the rapid prototyping and iterations that will be imperative to the success of the PipeSight platform’s Agile development process in Phase II. As part of the pipe stress analysis program, we also developed a custom visualizer that leverages the DOE-funded VTK visualization library. We added custom contouring capabilities and a means for interacting visually with both the inputs and outputs of the pipe stress analysis program. We also developed routines for automating the post-processing of the finite element simulations to determine if any failure criteria are met and to visualize the deformations, stresses and strains in ParaView using the exodus II file format (a subset of netCDF).

24 POWER TRANSMISSION AND DISTRIBUTION↗

A deep learning-guided automated workflow in LipidOz for detailed characterization of fungal fatty acid unsaturation by ozonolysis

Understanding fungal lipid biology and metabolism is critical for antifungal target discovery as lipids play central roles in cellular processes. Nuances in lipid structural differences can significantly impact their functions, making it necessary to characterize lipids in detail to enable and understanding of their roles in these complex systems. In particular, lipid double bond (DB) locations are an important component of lipid structure that can only be determined using a few specialized analytical techniques. Ozone-induced dissociation mass spectrometry (OzID-MS) is one such technique that uses ozone to break lipid DBs, producing pairs of characteristic fragments that allow the determination of DB positions. In this work we apply OzID-MS and LipidOz software to analyze the complex lipids of Saccharomyces cerevisiae yeast strains transfected with different fatty acid desaturases from Histoplasma capsulatum to determine the specific unsaturated lipids produce. The automated data analysis in LipidOz made the determination of DB positions from this large dataset more practical, but manual verification for all targets was still time-consuming. The DL model reduces manual involvement in data analysis, but since it was trained using mammalian lipid extracts, the prediction accuracy on yeast-derived data was reduced. We addressed both shortcomings by retraining the DL model to act as a pre-filter to prioritize targets for automated analysis, providing confident manually verified results but requiring less computational time and manual effort. Our workflow resulted in the determination of novel DB positions and enzymatic specificity.

mass spectrometry, deep learning, Lipidomics, doub↗

Low-Temperature Geothermal Resources: Relevant Data and PFA Methods to Reduce Development Risk: Preprint

This project is part of a larger national effort focused on demonstrating the multi-faceted value of integrating low-temperature geothermal resources into national decarbonization strategies and community energy plans. Low-temperature geothermal resources are defined as reservoirs-natural or engineered-with temperatures < 150 degrees C. While the focus in the NREL effort is on geothermal heating and cooling (GHC), resources at the upper end of this temperature range can also be used for small-scale power generation. However, low-temperature geothermal resources have not been studied as extensively as higher-temperature geothermal resources. We identified three major classes of low-temperature geothermal play types: sedimentary basins, orogenic systems, and radiogenic systems. We developed workflows for evaluating the potential of these resources building off the Play Fairway Analysis (PFA) approach to de-risking geothermal exploration. This PFA-based approach to low-temperature geothermal resources includes: (1) identifying relevant data; (2) grouping and weighting of relevant datasets into PFA criteria (e.g., geological, risk, economic criteria); (3) developing favorability or common risk maps for low-temperature geothermal resources to identify potential locations for more focused data collection; and (4) estimating electric power generation and heating potential at those locations using the GeoRePORT Resource Size Assessment Tool. This project will facilitate future deployment of GHC by providing data, tools, and workflows applicable to low-temperature geothermal resources.

favorability maps↗

Low-Temperature Geothermal Resources: Relevant Data and PFA Methods to Reduce Development Risk

This project is part of a larger national effort focused on demonstrating the multi-faceted value of integrating low-temperature geothermal resources into national decarbonization strategies and community energy plans. Low-temperature geothermal resources are defined as reservoirs-natural or engineered-with temperatures < 150 degrees C. While the focus in the NREL effort is on geothermal heating and cooling (GHC), resources at the upper end of this temperature range can also be used for small-scale power generation. However, low-temperature geothermal resources have not been studied as extensively as higher-temperature geothermal resources. We identified three major classes of low-temperature geothermal play types: sedimentary basins, orogenic systems, and radiogenic systems. We developed workflows for evaluating the potential of these resources building off the Play Fairway Analysis (PFA) approach to de-risking geothermal exploration. This PFA-based approach to low-temperature geothermal resources includes: (1) identifying relevant data; (2) grouping and weighting of relevant datasets into PFA criteria (e.g., geological, risk, economic criteria); (3) developing favorability or common risk maps for low-temperature geothermal resources to identify potential locations for more focused data collection; and (4) estimating electric power generation and heating potential at those locations using the GeoRePORT Resource Size Assessment Tool. This project will facilitate future deployment of GHC by providing data, tools, and workflows applicable to low-temperature geothermal resources.

favorability maps↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Fluorescent amplification for next generation sequencing (FA-NGS) library preparation

BACKGROUND: Next generation sequencing (NGS) has become a universal practice in modern molecular biology. As the throughput of sequencing experiments increases, the preparation of conventional multiplexed libraries becomes more labor intensive. Conventional library preparation typically requires quality control (QC) testing for individual libraries such as amplification success evaluation and quantification, none of which occur until the end of the library preparation process. RESULTS: In this study, we address the need for a more streamlined high-throughput NGS workflow by tethering real-time quantitative PCR (qPCR) to conventional workflows to save time and implement single tube and single reagent QC. We modified two distinct library preparation workflows by replacing PCR and quantification with qPCR using SYBR Green I. qPCR enabled individual library quantification for pooling in a single tube without the need for additional reagents. Additionally, a melting curve analysis was implemented as an intermediate QC test to confirm successful amplification. Sequencing analysis showed comparable percent reads for each indexed library, demonstrating that pooling calculations based on qPCR allow for an even representation of sequencing reads. To aid the modified workflow, a software toolkit was developed and used to generate pooling instructions and analyze qPCR and melting curve data. CONCLUSIONS: We successfully applied fluorescent amplification for next generation sequencing (FA-NGS) library preparation to both plasmids and bacterial genomes. As a result of using qPCR for quantification and proceeding directly to library pooling, the modified library preparation workflow has fewer overall steps. Therefore, we speculate that the FA-NGS workflow has less risk of user error. The melting curve analysis provides the necessary QC test to identify and troubleshoot library failures prior to sequencing. While this study demonstrates the value of FA-NGS for plasmid or gDNA libraries, we speculate that its versatility could lead to successful application across other library types.

59 BASIC BIOLOGICAL SCIENCES↗

Cell-Type-Specific Proteomics Analysis of a Small Number of Plant Cells by Integrating Laser Capture Microdissection with a Nanodroplet Sample Processing Platform

Plant organs and tissues contain multiple cell types, which are well organized in 3-dimensional structure to efficiently perform physiological functions such as homeostasis, response to environmental perturbation, pathogen infection. It is critically important to perform molecular measurements at the cell-type-specific level to discover mechanisms and unique features of cell populations that govern differentiation and respond to external perturbations. Although mass spectrometry-based proteomics has been demonstrated as an enabling discovery tool to study plant physiology, conventional approaches require millions of cells to generate robust biological conclusions. Such requirements mask the cell-to-cell heterogeneities and limit the comprehensive profiling of plant proteins at spatially resolved and cell-type-specific resolutions. This protocol describes a recently-developed proteomics workflow for studying a small number of plant cells by integrating laser capture microdissection, microfluidic nanodroplet-based sample preparation, with ultrasensitive liquid chromatography-mass spectrometry. Using poplar as a model tree species, we provide detailed protocols, including plant tissue harvest, tissue preparation, cryosectioning, laser microdissection, protein digestion, mass spectrometry measurement, and data analysis. We show the workflow enables the precise identification and quantification of thousands of proteins from hundreds of isolated plant root and leaf cells.

59 BASIC BIOLOGICAL SCIENCES↗