Python UQ Workflows Progress [Poster]
Project Overview and Goal: Create a Python Toolkit for UQ (PyTUQ) to facilitate the interoperability of FASTMath UQ tools (Dakota, UQTk, KLPC), as well as other ASCR funded capabilities for UQ workflows.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Project Overview and Goal: Create a Python Toolkit for UQ (PyTUQ) to facilitate the interoperability of FASTMath UQ tools (Dakota, UQTk, KLPC), as well as other ASCR funded capabilities for UQ workflows.
EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.
In this study, we used deep learning techniques, which are a form of artificial intelligence, to create fast and effective models for predicting how fluids flow in underground geological formations. This is important for managing geological carbon storage, a method used to fight climate change by storing carbon dioxide underground. The challenge lies in the complex nature of these underground spaces and the large amount of data needed to accurately simulate them. To overcome these issues, we developed a new workflow that reduces the data’s complexity before training the deep learning model and then reconstructs the predicted results in their original form. We also proposed a unique approach to handle the specific complexities found in 3D saturation fields, a crucial aspect of fluid flow prediction. We tested our method using real-world data from the Gulf of Mexico. Our results show that our approach not only accurately predicts fluid behavior but also significantly reduces computation time. This will greatly improve real-time decision-making and risk assessment in large-scale geological carbon storage operations.
The purpose of this project was to continue supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based protein and metabolite characterization. PNNL worked with Agilent to design, implement, evaluate, and demonstrate new algorithms and integrated them as functionalities into the PNNL-PreProcessor software. The project augmented PNNL’s capabilities to analyze complex proteomics and metabolomics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze these compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining liquid chromatography and ion mobility with mass spectrometry for widespread applications in life sciences and other fields.
NREL will provide technical assistance to help digitize best practice operations and maintenance (“O&M”) procedure guidelines utilizing the Illu platform for specific distributed solar-storage system configurations. NREL shall test Illu workflows in a real inspection environment for clarity of instruction and ease of documentation.
The poster describes the justIN workflow management system funded by the UK for DUNE and now used for all DUNE centrally managed data processing and simulation
Growing demands to understand PV deployment and reliability across an expanding range of climates, along with increasing computational power, are driving the need for simplified computing tools that support large-scale analysis and geospatial workflows. While existing open-source libraries and PV system modeling tools offer extensive collections of empirical and physical models, they often lack the ability to scale to large geospatial datasets. The tool presented here, GeoGridFusion, enables PV modelers to store and utilize gridded geospatial data from sources such as NSRDB, PVGIS, and others. GeoGridFusion harmonizes diverse datasets, taxonomies, and nomenclatures, and includes utilities that support intuitive geospatial area selection.
We explore how visualizations can help users understand what an AI agent is doing as it builds and runs queries over data. As part of the LinkQ system, a natural language interface for querying knowledge graphs with a large language model (LLM), we designed two complementary views: A State Diagram that shows where the agent is within a larger workflow, and a Live Action Display that gives real-time updates about the agent's current task. In a study with 14 practitioners, we found that these visuals helped participants build stronger mental models of the agent's behavior while also increasing their confidence in the system. However, we also observed that users sometimes trusted incorrect outputs simply because the agent appeared to be doing the "right" thing. Our findings point to both the value and risk of visualizing agent behavior in interactive AI systems.
Tutorials for modeling of slot-die and slide-die coating flows with Goma 7, an open source finite element code, are presented. The tutorials cover the workflow to attaining steady state solutions for these flows, and continuation strategies for navigating the operating windows. Advanced topics of coating window prediction, automated multiparameter continuation, non-Newtonian rheology, dynamic contact line modeling, and some more solution strategies are also covered.
The goal of the experiment was to demonstrate that the optimized multiplexed multi-PTM profiling workflow can comprehensively and quantitatively capture dynamic changes in protein abundance, cysteine oxidation, phosphorylation, and acetylation in cytokine-induced inflammatory stress in mouse pancreatic ß-cells. Global proteomic, redox proteomic, phosphoproteomic, and acetylomic were data collected from mouse Beta-TC-6 pancreatic Beta-cells, untreated (mock) and cytokine-treated Beta-cells at 4, 8, and 24 hours with 4 biological replicates. Samples were digested with trypsin and Lys-C, then analyzed by LC-MS/MS. Data were searched with MS-GF+, MASIC, and MaxQuant using PNNL's DMS processing pipeline.
Microeukaryotes (protists) serve fundamental roles in the marine environment as contributors to biogeochemical nutrient cycling and ecosystem function. Their activities can be inferred through metatranscriptomic investigations, which provide a detailed view into cellular processes, chemical-biological interactions in the environment, and ecological relationships among taxonomic groups. Established workflows have been individually put forth describing biomass collection at sea, laboratory RNA extraction protocols, and bioinformatic processing and computational approaches. Here, we present a compilation of current practices and lessons learned in carrying out metatranscriptomics of marine pelagic protistan communities, highlighting effective strategies and tools used by practitioners over the past decade. We anticipate that these guidelines will serve as a roadmap for new marine scientists beginning in the realms of molecular biology and/or bioinformatics, and will equip readers with foundational principles needed to delve into protistan metatranscriptomics.
Scientific computing heavily relies on data shared by the community, especially in distributed data-intensive applications. This research focuses on predicting slow connections that create bottlenecks in distributed workflows. In this study, we analyze network traffic logs collected between January 2021 and August 2022 at the National Energy Research Scientific Computing Center (NERSC). Based on the observed patterns, we define a set of features primarily based on history for identifying low-performing data transfers. Typically, there are far fewer slow connections on well-maintained networks, which creates difficulty in learning to identify these abnormally slow connections from the normal ones. We devise several stratified sampling techniques to address the class-imbalance challenge and study how they affect the machine learning approaches. Our tests show that a relatively simple technique that undersamples the normal cases to balance the number of samples in two classes (normal and slow) is very effective for model training. This model predicts slow connections with an F1 score of 0.926.
This workflow enables lamella production targeting fluorescently labeled biological structures that are small (<1 μm in axial extent) and rare (1 copy per cell) using a cryogenic tri-coincident imaging platform. In conclusion, this platform integrates fluorescence microscopy, focused ion beam milling, and scanning electron microscopy at a single focal position and enables simultaneous fluorescence microscopy while milling.
The Broadband Automation for Distributed Grid Efficiency and Resilience (BADGER) project aligns with national strategic priorities for integrating emerging wireless technologies and advancing AI-driven security. As critical infrastructure modernizes toward increasingly software-defined and interconnected systems, the ability to leverage 5G/NextG networks and AI-enabled control becomes essential. This report outlines work at the National Laboratory of the Rockies (NLR) to develop a NextG-native security architecture powered by AI-RAN concepts and evaluate workflows that enable efficient and reliable architectures. Together, these efforts position the laboratory to accelerate innovation while directly supporting national security and resilience objectives.
Microbes play fundamental roles in shaping natural ecosystem properties and functions, but do so under constraints imposed by their viral predators. However, studying viruses in nature can be challenging due to low biomass and the lack of universal gene markers. Though metagenomic short-read sequencing has greatly improved our virus ecology toolkit—and revealed many critical ecosystem roles for viruses—microdiverse populations and fine-scale genomic traits are missed. Some of these microdiverse populations are abundant and the missed regions may be of interest for identifying selection pressures that underpin evolutionary constraints associated with hosts and environments. Though long-read sequencing promises complete virus genomes on single reads, it currently suffers from high DNA requirements and sequencing errors that limit accurate gene prediction. Here we introduce VirION2, an integrated short- and long-read metagenomic wet-lab and informatics pipeline that updates our previous method (VirION) to further enhance the utility of long-read viral metagenomics. Using a viral mock community, we first optimized laboratory protocols (polymerase choice, DNA shearing size, PCR cycling) to enable 76% longer reads (now median length of 6,965 bp) from 100-fold less input DNA (now 1 nanogram). Using a virome from a natural seawater sample, we compared viromes generated with VirION2 against other library preparation options (unamplified, original VirION, and short-read), and optimized downstream informatics for improved long-read error correction and assembly. VirION2 assemblies combined with short-read based data (‘enhanced’ viromes), provided significant improvements over VirION libraries in the recovery of longer and more complete viral genomes, and our optimized error-correction strategy using long- and short-read data achieved 99.97% accuracy. In the seawater virome, VirION2 assemblies captured 5,161 viral populations (including all of the virus populations observed in the other assemblies), 30% of which were uniquely assembled through inclusion of long-reads, and 22% of the top 10% most abundant virus populations derived from assembly of long-reads. Viral populations unique to VirION2 assemblies had significantly higher microdiversity means, which may explain why short-read virome approaches failed to capture them. These findings suggest the VirION2 sample prep and workflow can help researchers better investigate the virosphere, even from challenging low-biomass samples. Our new protocols are available to the research community on protocols.io as a ‘living document’ to facilitate dissemination of updates to keep pace with the rapid evolution of long-read sequencing technology.
<span style="font-family: Calibri,sans-serif; font-size: 11pt;">Presentation on Nodeworks for the Workflow Summer Webinar Series (WoWoHa) 2020 given on August 28, 2020. See </span><a style="color: black; font-family: Calibri,sans-serif; font-size: 12pt;" href="https://wowoha.org/" rel="noopener noreferrer">https://wowoha.org/</a>
MOPED is an alternate workflow to HERON. Preliminary tests display computational advantages of MOPED in terms of run time. While maintaining these advantages, solution convergence remained similar.
Reduced chemistry models mitigate computational cost but introduce two sources of uncertainties in reacting flow simulation, including chemical information loss due to model reduction, and approximation errors due to non-optimal projection. We present an iterative workflow for quantification and minimization of reduced chemistry-induced uncertainties in reacting flow simulations.