CarrierSeq: a sequence analysis workflow for low-input nanopore sequencing
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The OSSE software provides an integrated end-to-end environment to simulate an Earth observing system by iteratively running a distributed modeling workflow based on the HyspIRI Mission, including atmospheric radiative transfer, surface albedo effects, detection, and retrieval for agile exploration of the mission design space. The software enables an Observing System Simulation Experiment (OSSE) and can be used for design trade space exploration of science return for proposed instruments by modeling the whole ground truth, sensing, and retrieval chain and to assess retrieval accuracy for a particular instrument and algorithm design. The OSSE in fra struc ture is extensible to future National Research Council (NRC) Decadal Survey concept missions where integrated modeling can improve the fidelity of coupled science and engineering analyses for systematic analysis and science return studies. This software has a distributed architecture that gives it a distinct advantage over other similar efforts. The workflow modeling components are typically legacy computer programs implemented in a variety of programming languages, including MATLAB, Excel, and FORTRAN. Integration of these diverse components is difficult and time-consuming. In order to hide this complexity, each modeling component is wrapped as a Web Service, and each component is able to pass analysis parameterizations, such as reflectance or radiance spectra, on to the next component downstream in the service workflow chain. In this way, the interface to each modeling component becomes uniform and the entire end-to-end workflow can be run using any existing or custom workflow processing engine. The architecture lets users extend workflows as new modeling components become available, chain together the components using any existing or custom workflow processing engine, and distribute them across any Internet-accessible Web Service endpoints. The workflow components can be hosted on any Internet-accessible machine. This has the advantages that the computations can be distributed to make best use of the available computing resources, and each workflow component can be hosted and maintained by their respective domain experts.
NASA ensures safe operation of complex systems through the use of formally-documented procedures, which encode the operational knowledge of the system as derived from system experts. Crew members use procedure documentation on the ground for training purposes and on-board space shuttle and space station to guide their activities. Investigators at JSC are developing a new representation for procedures that is content-based (as opposed to display-based). Instead of specifying how a procedure should look on the printed page, the content-based representation will identify the components of a procedure and (more importantly) how the components are related (e.g., how the activities within a procedure are sequenced; what resources need to be available for each activity). This approach will allow different sets of rules to be created for displaying procedures on a computer screen, on a hand-held personal digital assistant (PDA), verbally, or on a printed page, and will also allow intelligent reasoning processes to automatically interpret and use procedure definitions. During his NASA fellowship, Dr. Simpson examined how various industries represent procedures (also called business processes or workflows), in areas such as manufacturing, accounting, shipping, or customer service. A useful method for designing and evaluating workflow representation languages is by determining their ability to encode various workflow patterns, which depict abstract relationships between the components of a procedure removed from the context of a specific procedure or industry. Investigators have used this type of analysis to evaluate how well-suited existing workflow representation languages are for various industries based on the workflow patterns that commonly arise across industry-specific procedures. Based on this type of analysis, it is already clear that existing workflow representations capture discrete flow of control (i.e., when one activity should start and stop based on when other activities start and stop), but do not capture the flow of data, materials, resources or priorities. Existing workflow representation languages are also limited to representing sequences of discrete activities, and cannot encode procedures involving continuous flow of information or materials between activities.
Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.
NASA Earth Exchange (NEX), and her public cloud version OpenNEX, have become platforms supporting scientific collaboration, knowledge sharing and research for the entire Earth science community. To date, a number of custom tools and capabilities have been integrated into the platforms. However, such integration has to undergo a case-by-case manual process thus lacks scalability. This timely project builds an App Store onto OpenNEX as a building block. Climate data analytics tools/programs can be easily uploaded, shared, organized, searched, and recommended like photos and videos on the YouTube. The foundation of our App Store is a provenance server, which not only records metadata but also execution history of climate data analytics apps including the input data and parameters, output data and products, who runs the app for which purpose, and how apps may be chained into workflows. Researchers can thus understand, reproduce, and repurpose existing apps and workflows. Machine learning approaches are applied to mine provenance to provide recommend-as-you-go services for Earth scientists, such as to recommend suitable apps and workflow snippets. A browser-based workflow tool is also provided for researchers to explore the provenance server and design value-added workflows. Scalability, sustainability, extensibility, usability, adaptability, security and privacy are considered in the App Store.
As the overall manager and integrator of International Space Station (ISS) science payloads and experiments, the Payload Operations Integration Center (POIC) at Marshall Space Flight Center had a critical need to provide an information management system for exchange and management of ISS payload files as well as to coordinate ISS payload related operational changes. The POIC's information management system has a fundamental requirement to provide secure operational access not only to users physically located at the POIC, but also to provide collaborative access to remote experimenters and International Partners. The Payload Information Management System (PIMS) is a ground based electronic document configuration management and workflow system that was built to service that need. Functionally, PIMS provides the following document management related capabilities: 1. File access control, storage and retrieval from a central repository vault. 2. Collect supplemental data about files in the vault. 3. File exchange with a PMS GUI client, or any FTP connection. 4. Files placement into an FTP accessible dropbox for pickup by interfacing facilities, included files transmitted for spacecraft uplink. 5. Transmission of email messages to users notifying them of new version availability. 6. Polling of intermediate facility dropboxes for files that will automatically be processed by PIMS. 7. Provide an API that allows other POIC applications to access PIMS information. Functionally, PIMS provides the following Change Request processing capabilities: 1. Ability to create, view, manipulate, and query information about Operations Change Requests (OCRs). 2. Provides an adaptable workflow approval of OCRs with routing through developers, facility leads, POIC leads, reviewers, and implementers. Email messages can be sent to users either involving them in the workflow process or simply notifying them of OCR approval progress. All PIMS document management and OCR workflow controls are coordinated through and routed to individual user's "to do" list tasks. A user is given a task when it is their turn to perform some action relating to the approval of the Document or OCR. The user's available actions are restricted to only functions available for the assigned task. Certain actions, such as review or action implementation by non-PIMS users, can also be coordinated through automated emails.
To allow scientists further capabilities in the area of data mining and web services, the Goddard Earth Sciences Data and Information Services Center (GES DISC) and researchers at the University of Alabama in Huntsville (UAH) have developed a system to mine data at the source without the need of network transfers. The system has been constructed by linking together several pre-existing technologies: the Simple Scalable Script-based Science Processor for Measurements (S4PM), a processing engine at he GES DISC; the Algorithm Development and Mining (ADaM) system, a data mining toolkit from UAH that can be configured in a variety of ways to create customized mining processes; ActiveBPEL, a workflow execution engine based on BPEL (Business Process Execution Language); XBaya, a graphical workflow composer; and the EOS Clearinghouse (ECHO). XBaya is used to construct an analysis workflow at UAH using ADam components, which are also installed remotely at the GES DISC, wrapped as Web Services. The S4PM processing engine searches ECHO for data using space-time criteria, staging them to cache, allowing the ActiveBPEL engine to remotely orchestras the processing workflow within S4PM. As mining is completed, the output is placed in an FTP holding area for the end user. The goals are to give users control over the data they want to process, while mining data at the data source using the server's resources rather than transferring the full volume over the internet. These diverse technologies have been infused into a functioning, distributed system with only minor changes to the underlying technologies. The key to the infusion is the loosely coupled, Web-Services based architecture: All of the participating components are accessible (one way or another) through (Simple Object Access Protocol) SOAP-based Web Services.
The goal of NASA's Earthdata Search End-to-End Services workflow is to take the pain and headache out of searching for data and getting that data back in a format that is usable with only that data that is relevant for you. For too long scientists have had to jump through endless hoops, use tools that only offer specific data or specific services, and perform any number of other non-science tasks just to get started on their actual project. Earthdata Search leverages the Common Metadata Repository's (CMR) newly implemented Unified Metadata Models for Services and Variables as well as a new service broker to expose and seamlessly integrate a collection's service capabilities and variables into an intuitive user interface. Using the new End-to-End Services workflow, scientists will be able to quickly see what data is available to be customized, what customization options are available, and actually perform those customizations on the data all within Earthdata Search, regardless of who the data provider is. This talk will demonstrate the simple workflow that will be available to end users and also give an overview covering how the workflow is enabled by the metadata stored within the CMR. (https://search.earthdata.nasa.gov/)
The quantification and control of discretization error is critical to obtaining reliable simulation results. Adaptive mesh techniques have the potential to automate discretization error control, but have made limited impact on production analysis workflow. Recent progress has matured a number of independent implementations of flow solvers, error estimation methods, and anisotropic mesh adaptation mechanics. However, the poor integration of initial mesh generation and adaptive mesh mechanics to typical sources of geometry has hindered adoption of adaptive mesh techniques, where these geometries are often created in Mechanical Computer- Aided Design (MCAD) systems. The difficulty of this coupling is compounded by two factors: the inherent complexity of the model (e.g., large range of scales, bodies in proximity, details not required for analysis) and unintended geometry construction artifacts (e.g., translation, uneven parameterization, degeneracy, self-intersection, sliver faces, gaps, large tolerances be- tween topological elements, local high curvature to enforce continuity). Manual preparation of geometry is commonly employed to enable fixed-grid and adaptive-grid workflows by reducing the severity and negative impacts of these construction artifacts, but manual process interaction inhibits workflow automation. Techniques to permit the use of complex geometry models and reduce the impact of geometry construction artifacts on unstructured grid workflows are models from the AIAA Sonic Boom and High Lift Prediction are shown to demonstrate the utility of the current approach.
The Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni) is an online tool developed by the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers (DAACs) to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several well-known NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). There are twenty-two plots provided by Giovanni that can be used to analyze, compare, and explore Earth data across disciplines. Results can be shared with colleagues and downloaded for further analysis. Giovanni has helped publish over 3000 referral papers over the years. As open science policies roll in, data integrity has become a major challenge for Giovanni and other tools. For integrity, both data and workflows must be transparent. FAIR-compliant data, including input, intermediate, and result products, as well as their associated statistics, metadata, and information, are needed. The NASA Data Product Development Guide for Data Producers provides a key resource on how to develop FAIR-compliant data products. Data quality information is also needed from data producers and analysis services like Giovanni. The workflow part is quite challenging and requires workflow management improvements, such as recording workflows and making them available to users. In this presentation, we will discuss the data integrity challenges in Giovanni.
For the flooding seasons of 2011-2012 multiple space assets were used in a "sensorweb" to track major flooding in Thailand. Worldview-2 multispectral data was used in this effort and provided extremely high spatial resolution (2m / pixel) multispectral (8 bands at 0.45-1.05 micrometer spectra) data from which mostly automated workflows derived surface water extent and volumetric water information for use by a range of NGO and national authorities. We first describe how Worldview-2 and its data was integrated into the overall flood tracking sensorweb. We next describe the use of Support Vector Machine learning techniques that were used to derive surface water extent classifiers. Then we describe the fusion of surface water extent and digital elevation map (DEM) data to derive volumetric water calculations. Finally we discuss key future work such as speeding up the workflows and automating the data registration process (the only portion of the workflow requiring human input).
INTRODUCTION As with most computational analyses, a tradeoff exists between problem complexity, resource availability and response accuracy when modeling radiation transport from the source to a detector. The largest amount of analyst time for setting up an analysis is often spent ensuring that any simplifications made have minimal impact on the results. The vehicle shield geometry of interest is typically simplified from the original CAD design in order to reduce computation time, but this simplification requires the analyst to "re-draw" the geometry with a limited set of volumes in order to accommodate a specific radiation transport software package. The resulting low-fidelity geometry model cannot be shared with or compared to other radiation transport software packages, and the process can be error prone with increased model complexity. The work presented here demonstrates the use of the DAGMC (Direct Accelerated Geometry for Monte Carlo) Toolkit from the University of Wisconsin, to model the impacts of several space radiation sources on a CAD drawing of the US Lab module. METHODS The DAGMC toolkit workflow begins with the export of an existing CAD geometry from the native CAD to the ACIS format. The ACIS format file is then cleaned using SpaceClaim to remove small holes and component overlaps. Metadata is then assigned to the cleaned geometry file using CUBIT/Trelis from csimsoft (Registered Trademark). The DAGMC plugin script removes duplicate shared surfaces, facets the geometry to a specified tolerance, and ensures that the faceted geometry is water tight. This step also writes the material and scoring information to a standard input file format that the analyst can alter as desired prior to running the radiation transport program. The scoring results can be transformed, via python script, into a 3D format that is viewable in a standard graphics program. RESULTS The CAD model of the US Lab module of the International Space Station, inclusive of all the racks and components, was simplified to remove holes and volume overlaps. Problematic features within the drawing were also removed or repaired to prevent runtime issues. The cleaned drawing was then run through the DAGMC workflow to prepare for analysis. Pilot tests modeling transport of 1GeV proton and 800MeV/A oxygen sources show that reasonable results are converged upon in an acceptable amount of overall computation time from drawing preparation to data analysis. The FLUKA radiation transport code will next be used to model both a GCR and a trapped radiation source. These results will then be compared with measurements that have been made by the radiation instrumentation deployed inside the US Lab module. DISCUSSION Early analyses have indicated that the DAGMC workflow is a promising toolkit for running vehicle geometries of interest to NASA through multiple radiation transport codes. In addition, recent work has shown that a realistic human phantom, provided via a subcontract with the University of Florida, can be placed inside any vehicle geometry for a combinatorial analysis. This added functionality gives the user the ability to score various parameters at the organ level, and the results can then be used as input for cancer risk models.
Data publication is an essential activity for all data archives. Each of NASA's twelve Distributed Active Archive Centers (DAACs) have established publication workflows which account for the heterogeneous suite of missions, instruments, data providers, and datasets managed within the Earth Observation System Data and Information System (EOSDIS) program. Some aspects of data publication vary across DAACs: workflows range from manual to automatic, terms used to describe publication elements differ, and systems used to publish and manage data vary. Despite these differences, the DAAC data publication processes are generally the same: obtain the data and related information from data providers, describe the data with metadata and documentation, and release the data for access by the user community. In order to improve consistency and reduce the time required to publish data, we have developed a cross-DAAC initiative called the Common Earthdata Publication Framework (Earthdata Pub). Earthdata Pub seeks to: standardize communications and interactions with data providers; identify and standardize common workflows and steps in the data publication process; and design/implement a front-end system with features that include a common web interface, email & status tracking, and common application programming interfaces (APIs) to communicate with various DAAC-specific software components (services and applications) on the back-end. We will present the latest updates on this effort's progress and future plans.
Starting in 2017, NASA’s Human Research Program (HRP) Exploration Medical Capability (ExMC) element began a systems engineering transition from traditional, document-centric development to model-centric development when defining its foundation medical systems. These foundation medical systems define a Concept of Operations (ConOps) and identify the generic requirements for a medical system based on assumptions about a generic crew and mission environments and guidance from NASA standards (e.g., Medical “Levels of Care”). By making the transition, ExMC intends to improve communication among stakeholders about foundation medical system requirements and content. In addition, this transition will enable ExMC to lower both development and crew treatment risks for future, mission-specific medical systems. ExMC followed a Model Based Systems Engineering (MBSE) paradigm when developing the foundation medical systems. A model-based approach provides several advantages over a traditional, document-centric approach. First, when Systems Engineers (SE) develop diagrams in a model using a standard modeling language, they produce information dense pictures that facilitate understanding much more efficiently with less room for misinterpretation than text. Second, due to the evolving nature of projects, documentation becomes out of date the minute it is published. This can result in people making decisions based on information that is no longer current, especially if they are referencing a locally-stored copy of a document. A model, on the other hand, is always up to date with the latest approved changes and information. It serves as a single point of truth. Third, a model-centric approach centralizes all important information in one place. Rather than having to flip through separate ConOps documents, design specifications, requirements specifications, and the like to coordinate information, a model captures the content in one, integrated spot. This integration makes tracing information from end-to-end easier with greater reliability. The ExMC Systems Engineering Lifecycle follows a well-defined process. ExMC Systems Engineers perform all major steps of the process, regardless of the development methodology. One of the first steps in the process is developing the ConOps that describes the operation of the system from the point of view of the users. It includes a list of the users and their needs, the goals of the medical system, key assumptions about the system, and definitions of the medical system’s operational environments. For this development effort, ExMC chose to replace the traditional text-based ConOps document with a model. While the decision to change the development workflow was not difficult, implementing the structural and organizational workflows were. It required showing ExMC’s users, most of whom are not Systems Engineers, how the information they require would be presented in the model and to gain their acceptance of this approach. This paper documents key lessons learned during the ConOps transformation by focusing on how the model represents information, the agile workflow used by SEs when developing the model and how it integrates into a project plan, how leadership influenced key users to accept the transformation, and how the users interact with the model information.
This paper presents a new workflow for comparing experimental pressure-sensitive paint (PSP) data to computational fluid dynamic (CFD) simulations by way of mapping data from corresponding grids utilizing interpolation methods. In addition to generating quantitative and qualitative point-to-point comparisons between PSP and CFD data, this workflow extracts sectional loading data from both grids and generates lineload comparison charts for corresponding PSP and CFD runs. Experimental PSP data presented in this paper were taken from a 2016 NASA Ames Research Center Unitary Plan Wind Tunnel 11- by 11-Foot Transonic WindTunnel Facility test of the NASA Space Launch System. CFD simulation data for comparison purposes were generated using the FUN3D code. Overall, interpolation onto PSP grids versus CFD grids yields comparable surface pressure fields. However, lineload comparisons are easier to make on the CFD grid-mapped data due to the grid topology and the current capabilities of the lineload analysis tools at NASA Langley Research Center. This workflow is written using contemporary software (Python, Tecplot, PyTecplot), is compatible with existing tools at NASA Langley, and is developed to be adaptable depending on the situation.
Within NASA’s highly competitive environment for funding, Human Centered Design (HCD) and cybernetics could provide advantages to proposers during the mission formulation phase. Opportunities are limited when it comes to funding new science missions. Proposers are challenged to make a compelling case about the scientific desirability, technical feasibility, and resource viability of their concepts. Organizations follow established processes for proposal development using teams that typically include scientists, engineers, and managers. These team members are highly experienced subject matter experts (SME) in their own disciplines, and can respond to requirements from the solicitation. However, they are typically not trained as designers and communicators. Their approach is rooted within NASA’s science and technology paradigm. How can we improve the proposal development process, refine workflow between team members, and deliver clear and appealing offerings to the stakeholders and evaluators? These questions have been addressed by today’s most innovative companies (e.g., Apple, Google, 3M, Dyson), where the design process is not limited simply to engineering and management, but involves an all-encompassing approach drawing from fields such as social sciences, design, and the arts. Like these commercial enterprises, NASA currently employs systems thinking and integrated design, but can benefit further by moving beyond its current practices, which are mostly driven by rigid engineering, technology, science, and project management considerations. At JPL’s Innovation Foundry and through the Solar System Mission Formulation Office, we broadened this paradigm by including HCD in the mission formulation workflow. Our goal was to create a proposal with improved clarity and appeal, thus helping our team to communicate its message and aid evaluators with their work. In this paper we provide examples and lessons learned from our recent proposal development effort using HCD. We discuss touch points where we infused non-linear designerly approaches and cybernetic circularity into the workflow. Implemented design topics include operational design for team building; process design throughout distinct phases of the proposal development and writing process; communication design for streamlined exchange of information within the team and to stakeholders; interaction design; graphic design; and creating boundary objects. While these approaches may feel new or foreign to SMEs and managers in the aerospace community, they produced significant benefits in this mission formulation effort. We will describe how such approaches can be used to broaden NASA’s technology-driven paradigm through design, thus creating an environment which fosters innovation, improved communication, and strategic advantage for proposers and their organizations.
The NASA GeneLab project (genelab.nasa.gov) seeks to accelerate space biology research through cataloging and democratizing omics data. Since raw omics data is largely inaccessible to non-bioinformaticians, GeneLab works with the scientific community to develop standard processing pipelines to generate and publish processed data. Unlike raw data, processed data has greater immediate value to a wide range of users with varying technical backgrounds and computational capabilities. Standardizing processing workflows is essential to match the pace of raw data generation, ensure reproducibility, and enable standardized processed data for comparison across datasets. Previously, GeneLab developed a standardized pipeline for processing RNAseq data, referred to as the ‘GeneLab RNAseq Consensus Pipeline (RCP)’, in collaboration with GeneLab’s Analysis Working Groups. The work presented here is a Nextflow implementation of GeneLab’s RCP that automates and accelerates data processing of RNASeq datasets hosted on GeneLab. In addition to the core data processing, the workflow also includes staging of GeneLab raw data and a robust verification and validation (V&V) program that runs after each processing step to identify errors in real-time, stop additional downstream computation, and preserve computational resources. The workflow, including the staging and V&V functionality, is open source for others to reuse and modify at https://github.com/nasa/GeneLab_Data_Processing/tree/master/RNAseq.