Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Description of FY25 Theory and Simulation Performance Target: Development of an integrated modeling framework for fusion reactor design and assessment

The urgency to deliver fusion power is growing now more than ever, with increasing pressure for both public programs and private companies to meet milestones timelines and overcome significant remaining technical challenges to ensure growth of a nascent fusion industry in time to meet rapidly growing clean energy demands. With incredible advancements in computation and years of investment in fusion model development and validation, integrated modeling is poised to fill a key role in accelerating the timeline to a fusion pilot plant (FPP). Future fusion pilot plants will operate in regimes far beyond current experience, and device design will rely on physics-based prediction and extrapolation. Many concepts will also rely on simulation to assess safety (shielding, tritium management, materials activation and lifetimes), economics and scalability before the decision to build. Importantly, integrated simulation can be used to reveal and solve the complexities of system integration that may otherwise not be apparent in physical components or models developed in isolation. New experimental test facilities that produce relevant conditions to validate and resolve key technical challenges for various subsystems (materials, blankets, fuel cycle, etc.) have been repeatedly called for by the fusion community but are not yet realized. Integrated modeling has an important role in identifying realistic load conditions (thermal, electromagnetic, plasma, neutron and photon loads, etc.) and defining the components and experiments for these test facilities in order to ensure meaningful validation that sufficiently reduces modeling uncertainties and technical risk for the full integrated reactor. The Fusion REactor Design and Assessment (FREDA) SciDAC project is building a component-based integrated modeling framework & data structure to enable self-consistent, multi-fidelity, iterative optimization workflows for the fusion reactor design process. FREDA aims to shorten the time to viable designs by providing a set of flexible workflows to support the various stages of the design process using an integrated model hierarchy, ranging from the simple analytic descriptions to the highest fidelity, theory-based plasma and engineering modeling developed by the fusion and fission communities. These tools are expected to be needed for timely support of FPP design in the milestone program and in the FIRE collaboratives. The plasma simulation backbone of FREDA is IPS-FASTRAN with newly developed coupled Core-Edge Pedestal-SOL (CESOL) workflows, which is being extended to the far-SOL region up to the plasma facing components. FREDA incorporates the FERMI engineering modeling suite and will enable self-consistent evaluation of the thermal shields, limiters, blanket, magnets, and other surrounding structures with predictions of temperatures, erosion, dpa, activation, tritium generation and transport, creep, corrosion, material degradation, etc. Parametric generation of 3D CAD enables rapid iteration of component geometry in response to plasma and loading specifications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The Synthetic Biology Open Language (SBOL) Version 3: Simplified Data Exchange for Bioengineering

The Synthetic Biology Open Language (SBOL) is a community-developed data standard that allows knowledge about biological designs to be captured using a machine-tractable, ontology-backed representation that is built using Semantic Web technologies. While early versions of SBOL focused only on the description of DNA-based components and their sub-components, SBOL can now be used to represent knowledge across multiple scales and throughout the entire synthetic biology workflow, from the specification of a single molecule or DNA fragment through to multicellular systems containing multiple interacting genetic circuits. The third major iteration of the SBOL standard, SBOL3, is an effort to streamline and simplify the underlying data model with a focus on real-world applications, based on experience from the deployment of SBOL in a variety of scientific and industrial settings. Here, we introduce the SBOL3 specification both in comparison to previous versions of SBOL and through practical examples of its use.

59 BASIC BIOLOGICAL SCIENCES↗

Scientific Core Library Stack (SCLS) v2026

SCLS (Scientific Core Library Stack) is an opinionated build and packaging system for scientific computing libraries developed at Lawrence Berkeley National Laboratory. It produces a coherent, reproducible stack of numerical libraries — including BLAS/LAPACK, MPI, sparse direct and iterative solvers, graph partitioners, and parallel I/O libraries (e.g., PETSc, SLEPc, HDF5, NetCDF, MUMPS, OpenBLAS) — that work together without manual repair by downstream scientific software. From a single recipe-and-flavor model, SCLS produces native RPM packages for RHEL-family Linux, DEB packages for Debian/Ubuntu, direct Unix-style prefix installs for HPC and locked-down environments, and native macOS builds. Multiple build "flavors" (e.g., GCC+OpenBLAS, GCC+MKL, Intel+MKL, debug) coexist in distinct prefixes on the same host. Compared to general-purpose meta-build frameworks, SCLS is deliberately curated rather than infinitely configurable. It enforces deterministic, audit-friendly behavior: explicit build dependencies, no silent feature autodetection, a clear open-source license policy, and rpath-based runtime linkage so installs integrate cleanly with standard package-manager workflows.

Messe, Christian [Lawrence Berkeley National Labor↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Enhancing Monte Carlo Workflows for Nuclear Reactor Analysis with Metamodel-Driven Modeling

Monte Carlo codes are essential components of many reactor physics simulation workflows as high-fidelity continuous-energy neutron transport solvers. Among Monte Carlo radiation transport codes, MCNP is particularly notable due to its diverse simulation capabilities, large user base, and long validation history. Despite being a powerful simulation tool, MCNP provides limited capabilities to allow automated execution, model transformation, or support for user-defined logic and abstractions that limit its compatibility with modern workflows. Here, to better integrate MCNP into a modern scientific workflow, we have developed an intuitive yet full-featured MCNP Application Program Interface (API) in Python, named MCNPy, which provides a specialized set of classes for MCNP input development. Moreover, to guarantee that our reading, writing, and modeling capabilities remain self-consistent (and to render the huge scope of the MCNP API manageable), we have adopted a strategy of model-driven software development in which a generalized model of the MCNP input format has been created. From this generalized model, or “metamodel,” problem-specific implementations such as an engine for input validation or a codebase for programmatic operations may be automatically generated. Since MCNPy primarily acts as a Python front-end to the underlying Java API that directly interfaces with the metamodel, it is intrinsically linked to the metamodel and thus remains maintainable. With MCNPy, users can programmatically read, write, and modify any syntactically valid MCNP input file regardless of its origin. These capabilities allow users to automate complicated tasks like design optimization and model translation for nuclear systems. As examples, this work demonstrates the use of MCNPy to find the critical radius of a plutonium sphere and to translate a 9000+ line MCNP input file into a corresponding OpenMC model.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Integration of seismic-pressure-petrophysics inversion of continuous active-seismic monitoring data for monitoring and quantifying CO 2 plume (Final Report)

The overall objective of this project is to develop and validate an integrated package of joint seismic-pressure-petrophysics inversion (jSPPI) of continuous active-source seismic monitoring dataset capable of providing real-time monitoring of CO 2 plume during geologic carbon sequestration (GCS). The three specific developments include: (a) the methodologies for fast seismic full waveform inversion of continuous active source seismic monitoring, (CASSM) datasets for simultaneously estimating velocity and attenuation, and with data assimilation; (b) joint Bayesian petrophysical inversion of seismic models and pressure data for providing and updating CO 2 saturation models; (c) the methods using multiple datasets including (Crainfield and Frio-II borehole) synthetic, laboratory, and field CASSM datasets. The outcomes of jSPPI include (a) a workflow for processing CASSM data, (b) Bayesian inversion algorithms using CASSM data and pressure response data, and (c) integration with data assimilation algorithms for continuously updating site-specific models used for prediction and reservoir management. The validation of joint FWI will be conducted using synthetic models based on the Cranfield and Frio experiments as well as field CASSM datasets collected as part of the Frio-II pilot injection. To quantify and map the mass and distribution of CO 2 (saturation), we will jointly invert velocity and attenuation measurements from the FWI with a Bayesian approach using a rock physics model for attenuation (e.g., White’s attenuation model with two selected patch sizes (White, 1976; Dutta and Seriff, 1979)). The Bayesian inversion will be applied to each time step in the CASSM survey in an updating scheme, which integrates with an ensemble of reservoir simulations at each step. A more complete experimental validation dataset will be collected as part of a mesoscale (2-3 m) gas-CO 2 injection experiment utilizing a higher frequency version of the CASSM system developed for laboratory studies; the integrated inversion will be demonstrated using this dataset which will provide both a dense geometry as well as more precise secondary confirmation measurements (e.g. saturation) typically not available in the field. The resulting real-time map of CO 2 saturation is able to provide a deeper scientific understanding of the complex, time-varying dynamics of subsurface fluid flow migration path as well as the rapid detection of CO 2 leakage hazards.

25 ENERGY STORAGE↗

LatticeAnalytics: Strut-Level Visualization and Inspection of Additively Manufactured Lattice Structures

Additive manufacturing (AM) is revolutionizing the production of custom components with complex internal geometries, essential for high-performance applications in diverse fields such as medicine and defense. These AM parts optimize strength while minimizing weight by utilizing internal lattice structures consisting of large quantities of small interconnected struts. However, the complexity of these structures, combined with the challenges of using X-ray Computed Tomography (XCT) data, makes validation of part reliability difficult. This ultimately inhibits the development of novel parts for our collaborating material scientists. Here, we introduce LatticeAnalytics, a novel framework specifically designed for visual inspection of defects in these lattice structures. Our framework offers an end-to-end solution that includes the data management of XCT scans, enables remote access for geographically dispersed teams through a web-based dashboard, and incorporates novel visualizations. Our analysis is facilitated by a coarse alignment between the lattice’s nominal model, a spatial graph, and the XCT data. We employ a simple VR-based approach for fast and rough alignment, followed by an offline registration and identification of the struts. With the nodes and struts aligned and identified in the volume, our framework allows querying of subvolumes containing a single strut at multiple resolutions. This avoids computation over the entire lattice and also allow for easy parallelization of down-stream computations, such as strut-specific metrics. To depict a fast overview of the strut quality, we introduce two innovative visual encodings, crucial for our collaborators’ research in creating novel AM parts: the Contour View and the Roughness Map, which depict critical geometrical and surface features of individual struts in standardized two 2D views. We evaluated the integrated system through expert interviews. The feedback confirms the framework’s practicality and its effectiveness in enhancing current inspection workflows. It solves major bottlenecks for our collaborators, ultimately helping them create novel parts with advanced properties.

Miao, Haichao [Lawrence Livermore National Laborat↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

A KBase Case Study on Genome-wide Transcriptomics and Plant Primary Metabolism in Sorghum

A better understanding of the genetic and metabolic mechanisms that confer stress resistance and tolerance in plants is key to engineering new crops through advanced breeding technologies. This requires a systems biology approach that builds on a genome-wide understanding of the regulation of gene expression, plant metabolism, physiology and growth. In this study, we examine the response to drought stress in Sorghum, as we leverage the tools for transcriptomics and plant metabolic modeling we have implemented at the U.S. Department of Energy Systems Biology Knowledgebase (KBase). KBase enables researchers worldwide to collaborate and advance research by allowing them to upload private or public data into the KBase Narrative Interface, empowering them to analyze this data using a rich, extensible array of computational and data-analytics tools, and allows them to securely share scientific workflows and conclusions. We demonstrate how to use the current RNA-seq tools in KBase, applicable to both plants and microbes, to assemble and quantify long transcripts and identify differentially expressed genes effectively. More specifically, we demonstrate the utility of the platform by identifying key genes that are differentially expressed during drought-stress in Sorghum bicolor, which is an important sustainable production crop plant. We then show how to use KBase tools to predict the membership of genes in metabolic pathways and examine expression data in the context of metabolic subsystems. We demonstrate the power of the platform by making the data analysis and interpretation available to the biologists in the reproducible, re-usable, point-and-click format of a KBase Narrative thus promoting FAIR (Findable, Accessible, Interoperable and Reusable) guiding principles for scientific data management and stewardship.

59 BASIC BIOLOGICAL SCIENCES↗

Management and International Sorption Model Collaboration (M4SF-23LL010302062-NEA-TDB)

This progress report (Level 4 Milestone Number M4SF-23LL010302062) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Crystalline International Collaborations Activity Number SF-23LL01030206. The activity is focused on our long-term commitment of engaging our partners in international nuclear waste repository research. This includes participation in the Nuclear Energy Agency Thermochemical Database (NEA-TDB) Project and development of methodologies for integrating US and international thermodynamic databases for use in SFWST Generic Disposal System Assessment (GDSA) efforts. A continuing focus for FY23 efforts has been to support the US participation in the NEA-TDB effort (Mavrik Zavarin replaced Cindy Atkins-Duffin on the NEA-TDB Management Board (MB) and Executive Group (EG)) and developing mechanisms for integration of NEA-TDB thermochemical data with LLNL’s SUPCRTNE thermodynamic database that supports the SFWST GDSA activities. This effort is coordinated with the Argillite work package SUPCRTNE database development efforts. The goal is to provide a downloadable database that will be hosted on a LLNL website which integrates NEA-TDB data into the LLNL SUPCRTNE database where appropriate. As part of our international activities, we continue our effort to integrate international sorption databases into L-SCIE (Zavarin et al., 2022b). We presented opportunities to include sorption in the next phase of NEA-TDB efforts at the April 2023 EG meeting in Paris. FY23 efforts focused on ensuring interoperable database development across multiple international database development activities. The overall goal is to produce an open source database that can be shared and integrated with multiple nuclear waste programs internationally and harness modern data science workflows and algorithms to incorporate these new approaches into reactive transport and performance assessment models. In collaboration with our Helmholtz Zentrum Dresden Rossendorf partners, we recently demonstrated the power of FAIR open source databases by fitting iron oxide (hydrous ferric oxide, goethite, hematite, and magnetite) protolysis constants to all available L-SCIE data. The results were submitted as a manuscript to J. Colloid Interface Science. This work will inform future metal sorption studies on a variety of iron oxides in order to discern the most appropriate acidity constants and surface complexation modeling constructs to account for pH-dependent mineral surface charge behavior. This work also explored automated surface complexation model development workflows in order to generate higher throughput model input files for a more facile incorporation into GDSA activities.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

Where Is the Provenance? Ethical Replicability and Reproducibility in GIScience and Its Critical Applications

As replicability and reproducibility (R&R) crises develop within emerging convergent inquiry, ethical use of provenance information is central to the establishment and preservation of trust in critical applications of GIScience and geospatial technologies. Today large volumes of geospatial data are generated at high velocity from satellite sensors and unmanned aircraft systems, citizen sensors, geolocation-based data services, global navigation satellite systems, and so on. The extensive use of these data for applications such as disaster and humanitarian response raises the issue of R&R from competing perspectives of location privacy and geospatial data quality. Although geospatial data can be integrated and linked with contextual information to identify individuals’ movements, steps taken to ensure privacy can complicate the multiuser development of high-quality geospatial workflows. Provenance information as digital records of historical (retrospective) and potential future (prospective) geospatial processes is often overlooked, misunderstood, or inadequately addressed. We explore the relationship between provenance information, location privacy, and geospatial data quality in the context of R&R with a focus on disaster analytics. Here, we argue that in the era of big data and deep learning, GIScientists and associated institutions bear greater responsibility both for geospatial workflow quality and for location privacy. Given vastly heterogenous computational landscapes, we provide practical recommendations for ethically driven provenance and R&R research and development within the GIScience community and beyond.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

November 2021 Operational Highlights

Completed: 1. the initial tests of document workflows for Titan on the Red, which included standing up a software encryption capability to test data ingestion; 2. a software encryption capability in order to ingest the first data source by the end of the month; 3. a stand-alone computer for classified ontology data entry. Titan on the Red is an artificial intelligence/machine learning system to make digitizing, cataloging, and searching NSRC collections easier and more efficient. Created Online Vault backups, with one copy stored at LANL and one shipped to Lawrence Livermore National Laboratory. The Online Vault is a classified, searchable library of LANL’s nuclear weapons design and test history. Imported new Laboratory Directed Research and Development (LDRD) documents based on revised access categories. Bulk ingested nearly 10,000 documents via java-based PowerLoader. This application ingests metadata and content into the Online Vault to meet the requirements for the NSRC collections. Rehoused 300 linear feet of weapons physics documents in archival, acid-free storage folders and boxes.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Geologic hydrogen: From natural occurrences to anthropogenic generation – A review of fundamentals, potential, challenges and prospects

Growing demand for hydrogen is exposing the environmental and economic limits of reforming-based and carbon-managed supply chains, while the scale-up of electrolytic capacity remains capital-constrained. Geologic hydrogen, defined as molecular H₂ generated and stored within the Earth's crust offers a complementary, potentially lower-cost resource, yet exploration is still ad hoc. This review (1) revisits a global inventory of confirmed hydrogen seeps and subsurface occurrences; (2) analyzes the controlling reactions, migration pathways, and trapping conditions governing these occurrences; (3) proposes a process-based geologic hydrogen system concept analogous to, yet distinct from, the petroleum system; and (4) evaluates potential geologic hydrogen systems within the United States as a representative case study. Here, we contrast natural systems powered by serpentinization, mantle degassing or radiolysis with anthropogenic systems that stimulate the same reactions or convert in-situ hydrocarbons. Stable hydrogen accumulations require generation rates that outpace combined physical, chemical and microbial losses; the Bourakébougou field (Mali) exemplifies a self-recharging, free-gas reservoir sustained by meteoric-water serpentinization beneath an efficient caprock. Prospective geologic hydrogen resources are likely to occur in regions where iron-rich lithologies, deep-seated faults, and low-permeability sealing formations coexist. Applying this principle, we highlight three promising hydrogen play types in U.S. geological terrains: ophiolite belts (Appalachian and Californian regions), the Midcontinent Rift and the Lake Superior banded‑iron formations. Multiphysics numerical models and positive-unlabeled machine-learning workflows help to accelerate play screening and de-risk future production; yet, reaction kinetics, stimulation strategies, and full techno-economic and life-cycle assessments remain pivotal knowledge gaps.

Anthropogenic hydrogen generation↗

Automatic building energy model development and debugging using large language models agentic workflow

Building energy modeling (BEM) is a complex process that demands significant time and expertise, limiting its broader application in building design and operations. While Large Language Models (LLMs) agentic workflow have facilitated complex engineering processes, their application in BEM has not been specifically explored. This paper investigates the feasibility of automating BEM using LLM agentic workflow. Here, we developed a generic LLM-planning-based workflow that takes a building description as input and generates an error-free EnergyPlus building energy model. Our robust workflow includes four core agents: 1) Building Description Pre-Processing, 2) IDF Object Information Extraction, 3) Single IDF Object Generator Suite, and 4) IDF Debugging Agent. These agents divide the complex tasks into manageable sub-steps, enabling LLMs to generate accurate and reliable results at each stage. The case study demonstrates the successful translation of a building description into an error-free EnergyPlus model for the iUnit modular building at the National Renewable Energy Laboratory. The effectiveness of our workflow surpasses: 1) naive prompt engineering, 2) other LLM-based workflows, and 3) manual modeling, in terms of accuracy, reliability, and time efficiency. The paper concludes with a discussion on the interplay between foundational models and LLM agent planning design, advocating for the use of fine-tuned, specialized models to advance this field.

97 MATHEMATICS AND COMPUTING↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Conference presentation at Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Extended abstract for Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗