Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Hierarchical Control and Stability Analysis for a Nonisolated Grid-Tied DC Energy Router Integrating Energy Storage and Partial Distributed Generation

This article proposes a nonisolated dc converter-based energy router (dc-ER) and its operating strategy. Here, the intent is to integrate energy storage (ES), distributed generation (DG), local load, and dc power grid in an autonomous and more efficient way. The ES/DG power ports are coupled with each other in partial-type connections, which obtains higher voltage supply gain and direct input–output power transmission. High-voltage supply gain allows a wide solar power tracking range and the possibility of optimal battery charging/discharging. Direct input–output power transmission improves the energy conversion efficiency. In this article, the operating modes of dc-ER are first analyzed, followed by the optimal hardware design counting the DG current ripple minimization and the limitation caused by the power flow direction. Second, the mathematical model of dc-ER is deduced. The hierarchical control is then proposed to manage port energy in a flexible manner. The stability is analyzed by using impedance modeling. Finally, experimental and simulation results verify the superiorities of the proposed topology in terms of flexible operating mode transition and high-power conversion efficiency.

25 ENERGY STORAGE↗

NASA/IEEE MSST 2004 Twelfth NASA Goddard Conference on Mass Storage Systems and Technologies in cooperation with the Twenty-First IEEE Conference on Mass Storage Systems and Technologies

MSST2004, the Twelfth NASA Goddard / Twenty-first IEEE Conference on Mass Storage Systems and Technologies has as its focus long-term stewardship of globally-distributed storage. The increasing prevalence of e-anything brought about by widespread use of applications based, among others, on the World Wide Web, has contributed to rapid growth of online data holdings. A study released by the School of Information Management and Systems at the University of California, Berkeley, estimates that over 5 exabytes of data was created in 2002. Almost 99 percent of this information originally appeared on magnetic media. The theme for MSST2004 is therefore both timely and appropriate. There have been many discussions about rapid technological obsolescence, incompatible formats and inadequate attention to the permanent preservation of knowledge committed to digital storage. Tutorial sessions at MSST2004 detail some of these concerns, and steps being taken to alleviate them. Over 30 papers deal with topics as diverse as performance, file systems, and stewardship and preservation. A number of short papers, extemporaneous presentations, and works in progress will detail current and relevant research on the MSST2004 theme.

Kobler, Ben↗

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

Platform for efficient large-scale storage and analysis of multi-omics data in plant and microbial systems (Final Technical Report)

Genomic variation at the sequence level fundamentally affects the phenotypic state of all organisms at all stages of development, while dynamic processes such as changes in the epigenome (e.g. DNA methylation state) and transcriptome regulate the specific phenotype expressed at any given state of development based upon that genomic variation. In plants, DNA methylation is a particularly important mechanism for both regulating transcriptomic expression and for management of genomic variations that could be deleterious to the organism due to the presence of active retrotransposons in plant genomes. While DNA methylation is heritable, it is also dynamic through a given plant’s development and life cycle, particularly during the development from seed to mature specimen suggesting variations in DNA methylation could be critical regulators of biologically and commercially important phenotypes such as time to flowering; in addition, plant DNA methylation is more complex than that of animals, with methylation of CHG and CHH trinucleotides evident in addition to the better-known CG methylation. The complexity of plant DNA methylation and its interplay with genomic sequence variation, transcriptomics and other epigenomic factors demand a storage and analysis framework that can cope with the complexity both within a single specimen and with analyses that span many individuals and even many species, such as attempts to extend models from model organisms to commercially relevant species. In addition to complexity, the rapid development and proliferation of sequencing technology has led to an explosion of data volume that conventional storage and analysis solutions will likely be unable to cope with in the long run. We proposed to study these with suitable distributed storage and computation and therefore for the application of cloud computing to biological analyses; integrate with existing data sources and compatible with virtually any interface use case, from fully automated shell scripts to notebooks and do all these at scale in this STTR grant.

60 APPLIED LIFE SCIENCES↗

Hierarchical Control of Distributed Battery Energy Storage Systems in a DC Microgrid

Contents - NASA Future Power Systems - DC Microgrid with Distributed Battery Energy Storage - A simplified gateway power system - Droop Control (Primary Control) - Distributed battery units, droop control, and load sharing characteristics - Simulation results - Three-Level Hierarchical Control of Distributed Battery Units During Eclipse - Secondary Control - Objective and definition of unit control error - Implementation of secondary control - Simulation results - Tertiary Control - Schedule BU current sharing with the objective of equal SoC - Simulation results•Conclusions

Jing Zhang↗

Redundant Disk Arrays in Transaction Processing Systems

We address various issues dealing with the use of disk arrays in transaction processing environments. We look at the problem of transaction undo recovery and propose a scheme for using the redundancy in disk arrays to support undo recovery. The scheme uses twin page storage for the parity information in the array. It speeds up transaction processing by eliminating the need for undo logging for most transactions. The use of redundant arrays of distributed disks to provide recovery from disasters as well as temporary site failures and disk crashes is also studied. We investigate the problem of assigning the sites of a distributed storage system to redundant arrays in such a way that a cost of maintaining the redundant parity information is minimized. Heuristic algorithms for solving the site partitioning problem are proposed and their performance is evaluated using simulation. We also develop a heuristic for which an upper bound on the deviation from the optimal solution can be established.

Mourad, Antoine Nagib↗

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING↗

Technology for national asset storage systems

An industry-led collaborative project, called the National Storage Laboratory, was organized to investigate technology for storage systems that will be the future repositories for our national information assets. Industry participants are IBM Federal Systems Company, Ampex Recording Systems Corporation, General Atomics DISCOS Division, IBM ADSTAR, Maximum Strategy Corporation, Network Systems Corporation, and Zitel Corporation. Industry members of the collaborative project are funding their own participation. Lawrence Livermore National Laboratory through its National Energy Research Supercomputer Center (NERSC) will participate in the project as the operational site and the provider of applications. The expected result is an evaluation of a high performance storage architecture assembled from commercially available hardware and software, with some software enhancements to meet the project's goals. It is anticipated that the integrated testbed system will represent a significant advance in the technology for distributed storage systems capable of handling gigabyte class files at gigabit-per-second data rates. The National Storage Laboratory was officially launched on 27 May 1992.

Coyne, Robert A.↗

Grid-Forming Storage Networks: Analytical Characterization of Damping and Design Insights

Grid-forming storage resources are critical to the operation of power grids with high renewable penetration. Over the years, there has been considerable emphasis on understanding the benefits these inverter-based storage resources bring towards managing volatility and enhancing grid reliability but little is studied about the supplementary advantages latent in their operation. This paper investigates one such application in which we explore the impact of grid-forming distributed storage in damping low-frequency inter-area oscillations. We present a detailed analysis characterizing the impact of inverter-droop and storage size on the slower eigenvalues, highlighting potential design considerations for enhancing system stability.

Chatterjee, Kaustav [BATTELLE (PACIFIC NW LAB)] (O↗

Strong and Efficient Consistency with Consistency-aware Durability

We introduce consistency-aware durability or C ad , a new approach to durability in distributed storage that enables strong consistency while delivering high performance. We demonstrate the efficacy of this approach by designing cross-client monotonic reads , a novel and strong consistency property that provides monotonic reads across failures and sessions in leader-based systems; such a property can be particularly beneficial in geo-distributed and edge-computing scenarios. We build O rca , a modified version of ZooKeeper that implements C ad and cross-client monotonic reads. We experimentally show that O rca provides strong consistency while closely matching the performance of weakly consistent ZooKeeper. Compared to strongly consistent ZooKeeper, O rca provides significantly higher throughput (1.8--3.3×) and notably reduces latency, sometimes by an order of magnitude in geo-distributed settings. We also implement C ad in Redis and show that the performance benefits are similar to that of C ad ’s implementation in ZooKeeper.

Computer Science↗

Safety of Mobile Hydrogen and Fuel Cell Technology Applications: An Investigation by the Hydrogen Safety Panel

Safe practices in the production, storage, distribution, and use of hydrogen are essential for the widespread acceptance of hydrogen and fuel cell technologies. A significant safety incident could damage public perception of hydrogen and fuel cells. Recent incidents involving multi-cylinder hydrogen transport vehicles in the United States have brought attention to the potential impacts of mobile hydrogen storage and transportation. Road transportation of bulk gaseous hydrogen presents unique hazards that can be very different from those for stationary equipment, and new equipment developers may have less experience and expertise than seasoned gas providers. In response to the aforementioned incidents, and in support of hydrogen and fuel cell activities in California specifically, the Hydrogen Safety Panel (HSP) has investigated the safety of mobile hydrogen and fuel cell applications (mobile auxiliary/emergency fuel cell power units, mobile fuelers, multi-cylinder transport vehicles, unmanned aircraft power supplies, and mobile hydrogen generators). The HSP examined the applications, requirements, and performance of mobile applications that are being used extensively outside of California to understand how safety considerations are applied. This report discusses the results of the HSP’s evaluation of hydrogen and fuel cell mobile applications along with recommendations to address relevant safety issues.

08 HYDROGEN↗

Benchmark Comparison of Cloud Analytics Methods Applied to Earth Observations

Earth Observation data are a vital resource for studying long term changes, but the large data volumes can be challenging to analyze. Time series analysis in particular is hampered by the typical thin-time-slice file organization. We examine several potential solutions inspired in large part by the data-parallel methods that have arisen with cloud computing. These solutions include various combinations of data re-organization, spatial indexing, distributed storage and pre-computation that we term "Analytics Optimized Data Stores" (AODS). We find that even simple solutions (such as a data cube) produce more than an order of magnitude improvement; the best provide two to three orders of magnitude improvement. The most performant solutions have tradeoffs in terms of generality or storage footprint, but may nonetheless be useful components in data analytics frameworks where performance is critical.

parallel processing (computers)↗

Data-driven offshore CO 2 saline storage assessment methodology

The world produces approximately 50 billion tonnes of greenhouse gases annually. This is measured in CO 2 -equivalents, and geologic CO 2 storage has the potential to advance decarbonization and mitigate greenhouse gas emissions. New technologies to assess offshore carbon storage are needed to address resource, regulatory, and commercial needs. Although most efforts to assess storage resources focus on onshore criteria, offshore reservoirs offer significant storage potential and distinct development challenges. Potential advantages of offshore carbon storage include being further from human population centers and less potential to interact with groundwater. The U.S. Department of Energy's method for evaluating storage capacity in non-oil-bearing saline reservoirs has been enhanced to support assessments for offshore environments in the Offshore CO 2 Saline Storage methodology (OCSS). This methodology applies data-driven capabilities to estimate saline storage capacity while accounting for features specific to offshore reservoirs. Features include changes in CO 2 density and sedimentary differences that impact estimates of permanence and capacity. The Offshore CO 2 Saline Storage Calculator mechanizes OCSS to estimate storage capacity. This paper presents the methodology and estimates for 18 geologic domains in the Gulf of Mexico. Potential storage distributions, sensitivity analyses, and the incorporation of spatial data and tools to support safe site selection are also discussed.

58 GEOSCIENCES↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Community-Scale Solar Deployment in the Northwest Arctic

NANA Regional Corporation (NRC, or NANA) was formed as an Alaska Native Corporation (ANC) pursuant to the Alaska Native Claims Settlement Act of 1971. Our lands cover 39,000 square miles of the Northwest Arctic region of Alaska. We partner closely with the Northwest Arctic Borough (NAB) and our 11 remote communities on numerous projects and activities, especially around clean energy initiatives that improve quality of life and help to reduce extremely high energy costs experienced by the communities in our region. Collectively, the regional partnership has developed numerous successful solar, wind, biomass, and energy storage, distribution upgrade, and efficiency projects. To further progress, we created the Northwest Arctic Energy Steering Committee to share and replicate these benefits across the region. NANA’s mission is to provide economic opportunities for our more than 13,500 Iñupiat shareholders and to protect and enhance NANA lands.

14 SOLAR ENERGY↗

Soil Carbon Stocks Not Linked to Aboveground Litter Input and Chemistry of Old-Growth Forest and Adjacent Prairie

The long-standing assumption that aboveground plant litter inputs have a substantial influence on soil organic carbon storage (SOC) and dynamics has been challenged by a new paradigm for SOC formation and persistence. We tested the importance of plant litter chemistry on SOC storage, distribution, composition, and age by comparing two highly contrasting ecosystems: an old-growth coast redwood (Sequoia sempervirens) forest, with highly aromatic litter, and an adjacent coastal prairie, with more easily decomposed litter. We hypothesized that if plant litter chemistry was the primary driver, redwood would store more and older SOC that was less microbially processed than prairie. Total soil carbon stocks to 110 cm depth were higher in prairie (35 kg C m —2 ) than redwood (28 kg C m —2 ). Radiocarbon values indicated shorter SOC residence times in redwood than prairie throughout the profile. Higher amounts of pyrogenic carbon and a higher degree of microbial processing of SOC appear to be instrumental for soil carbon storage and persistence in prairie, while differences in fine-root carbon inputs likely contribute to younger SOC in redwood. We conclude that at these sites fire residues, root inputs, and soil properties influence soil carbon dynamics to a greater degree than the properties of aboveground litter.

13C-NMR spectroscopy↗

Hierarchical Control of Distributed Battery Energy Storage System in a DC Microgrid

This paper presents a novel hierarchical control approach of a DC microgrid (DCMG) which is supplied by a distributed battery energy storage system (BESS). With this approach, all battery units distributed in the BESS can be controlled to discharge with accurate current sharing and state-of-charge (SoC) balancing. Similar to other hierarchical control approaches used in DCMGs, this approach consists of three levels: (1) primary control, (2) secondary control and (3) tertiary control. This work includes defining a unit control error (UCE) at the secondary control level and evaluating current sharing weights at tertiary control level. A centralized controller at secondary control level is designed to detect the UCEs of each battery unit, and to restore the average voltage of a DCMG and control battery current sharing simultaneously. The distributed battery units share the load current in a DCMG based on weights. These weights are evaluated at the tertiary control level based on battery SoCs. The approach’s effectiveness was confirmed in digital simulation tests with the same simulation model as used in the NASA AMPS Modular Hardware Emulator.

Jing Zhang↗

U.S. national water and energy land dataset for integrated multisector dynamics research

Abstract Understanding resource demands and tradeoffs among energy, water, and land socioeconomic sectors requires an explicit consideration of spatial scale. However, incorporation of land dynamics within the energy-water nexus has been limited due inconsistent spatial units of observation from disparate data sources. Herein we describe the development of a National Water and Energy Land Dataset (NWELD) for the conterminous United States. NWELD is a 30-m, 86-layer rasterized dataset depicting the land use of mappable components of the United States energy sector life cycles (and related water used for energy), specifically the extraction, development, production, storage, distribution, and operation of eight renewable and non-renewable technologies. Through geospatial processing and programming, the final products were assembled using four different methodologies, each depending upon the nature and availability of raw data sources. For validation, NWELD provided a relatively accurate portrayal of the spatial extent of energy life cycles yet displayed low measures of association with mainstream land cover and land use datasets, indicating the provision of new land use information for the energy-water nexus.

58 GEOSCIENCES↗