Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “storage throughput”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning-assisted design of metal–organic frameworks for hydrogen storage: A high-throughput screening and experimental approach

Various theoretical approaches, including big data and high-throughput screening techniques, have been explored in developing new materials due to their significant potential time-saving advantages. However, it remains a significant challenge to experimentally realize new materials that are predicted. In this study, we propose a novel materials design strategy that utilizes machine-learning (ML) techniques to predict new porous materials that show promise for hydrogen storage and are likely to be feasible to synthesize. By leveraging ML techniques and metal–organic framework (MOF) databases, we are able to predict the synthesizability of MOF structures. This is evidenced by the successful synthesis of a new vanadium-based MOF that exhibits excellent performance for cryogenic H 2 storage. Notably, the total gravimetric and volumetric H 2 uptakes are as high as 9.0 wt% and 50.0 g/L at 77 K and 150 bar. This ML-assisted materials design offers an efficient and promising approach for developing hydrogen storage materials.

08 HYDROGEN↗

Striped tape arrays

A growing number of applications require high capacity, high throughput tertiary storage systems. How data striping ideas apply to arrays of magnetic tape drives is investigated. Data striping increases throughput and reduces response time for large accesses to a storage system. Striped magnetic tape systems are particularly appealing because many inexpensive magnetic tape drives have low bandwidth; striping may offer dramatic performance improvements for these systems. There are several important issues in designing striped tape systems: the choice of tape drives and robots, whether to stripe within or between robots, and the choice of the best scheme for distributing data on cartridges. One of the most troublesome problems in striped tape arrays is the synchronization of transfers across tape drives. Another issue is how improved devices will affect the desirability of striping in the future. The results of simulations comparing the performance of striped tape systems to non-striped systems are presented.

Drapeau, Ann L.↗

Comparison of Classical and Charge Storage Methods for Determining Conductivity of Thin Film Insulators

Conductivity of insulating materials is a key parameter to determine how accumulated charge will distribute across the spacecraft and how rapidly charge imbalance will dissipate. Classical ASTM and IEC methods to measure thin film insulator conductivity apply a constant voltage to two electrodes around the sample and measure the resulting current for tens of minutes. However, conductivity is more appropriately measured for spacecraft charging applications as the "decay" of charge deposited on the surface of an insulator. Charge decay methods expose one side of the insulator in vacuum to sequences of charged particles, light, and plasma, with a metal electrode attached to the other side of the insulator. Data are obtained by capacitive coupling to measure both the resulting voltage on the open surface and emission of electrons from the exposed surface, as well monitoring currents to the electrode. Instrumentation for both classical and charge storage decay methods has been developed and tested at Jet Propulsion Laboratory (JPL) and at Utah State University (USU). Details of the apparatus, test methods and data analysis are given here. The JPL charge storage decay chamber is a first-generation instrument, designed to make detailed measurements on only three to five samples at a time. Because samples must typically be tested for over a month, a second-generation high sample throughput charge storage decay chamber was developed at USU with the capability of testing up to 32 samples simultaneously. Details are provided about the instrumentation to measure surface charge and current; for charge deposition apparatus and control; the sample holders to properly isolate the mounted samples; the sample carousel to rotate samples into place; the control of the sample environment including sample vacuum, ambient gas, and sample temperature; and the computer control and data acquisition systems. Measurements are compared here for a number of thin film insulators using both methods at both facilities. We have found that conductivity determined from charge storage decay methods is 102 to 104 larger than values obtained from classical methods. Another Spacecraft Charging Conference presentation describes more extensive measurements made with these apparatus. This work is supported through funding from the NASA Space Environments and Effects Program and the USU Space Dynamics Laboratory Enabling Technologies Program.

Swaminathan, Prasanna↗

When to use rsync

We have endeavored to show, using a series of data transfer results obtained from two testbeds, when to use the popular data copying tool rsync and related tools. Tests have been conducted in local area network (LAN) and wide area network (WAN) environments. We conclude that for files in a certain size range and network latency ≦ 10 ms round trip time (RTT), rsync is still useful for data moving tasks in the category 4 of the U.S. DOE Technical Report “Data Movement Categories”. For more demanding data movement requirements, tools of different classes are suggested. Sample histograms from two DOE user facilities are provided to further support our conclusions.

97 MATHEMATICS AND COMPUTING↗

The crush of new earth science data knocking at our door

The reasons for collecting massive amounts of earth science data in the Earth Observing System (EOS) Project are discussed. A processing hierarchy for handling the data is described, and the prospects for adequate throughput and storage for operational data analysis in the EOS era are addressed. Needs for successful exploratory data analysis are examined, and the policy issues implicated by the large stream of EOS data are considered.

Kahn, Ralph↗

dCache project status and update

The dCache project delivers an open-source, massively scalable, distributed storage system deployed internationally to satisfy today’s scientists’ ever-demanding storage requirements. Its multifaceted approach supports different use cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and longterm data persistence on tertiary storage. Even though dCache was initially developed for HEP experiments, today, it is used by various scientific communities, including astrophysics, biomed, and life science, each with their specific requirements. To match the needs of these new communities and keep up with the scaling demands of existing experiments, dCache is permanently evolving. With this contribution, we would like to highlight the recent developments in dCache regarding integration with CERN Tape Archive (CTA), advanced metadata handling, token-based authorization support, bulk API for QoS transitions, REST API to control interaction with the tape system, and future development directions.

Mkrtchyan, Tigran [DESY]↗

Performance evaluation of the Engineering Analysis and Data Systems (EADS) 2

The Engineering Analysis and Data System (EADS)II (1) was installed in March 1993 to provide high performance computing for science and engineering at Marshall Space Flight Center (MSFC). EADS II increased the computing capabilities over the existing EADS facility in the areas of throughput and mass storage. EADS II includes a Vector Processor Compute System (VPCS), a Virtual Memory Compute System (CFS), a Common Output System (COS), as well as Image Processing Station, Mini Super Computers, and Intelligent Workstations. These facilities are interconnected by a sophisticated network system. This work considers only the performance of the VPCS and the CFS. The VPCS is a Cray YMP. The CFS is implemented on an RS 6000 using the UniTree Mass Storage System. To better meet the science and engineering computing requirements, EADS II must be monitored, its performance analyzed, and appropriate modifications for performance improvement made. Implementing this approach requires tool(s) to assist in performance monitoring and analysis. In Spring 1994, PerfStat 2.0 was purchased to meet these needs for the VPCS and the CFS. PerfStat(2) is a set of tools that can be used to analyze both historical and real-time performance data. Its flexible design allows significant user customization. The user identifies what data is collected, how it is classified, and how it is displayed for evaluation. Both graphical and tabular displays are supported. The capability of the PerfStat tool was evaluated, appropriate modifications to EADS II to optimize throughput and enhance productivity were suggested and implemented, and the effects of these modifications on the systems performance were observed. In this paper, the PerfStat tool is described, then its use with EADS II is outlined briefly. Next, the evaluation of the VPCS, as well as the modifications made to the system are described. Finally, conclusions are drawn and recommendations for future worked are outlined.

Debrunner, Linda S.↗

Variable rate neural compression for sparse detector data

Particle colliders produce data at extraordinary rates, posing major challenges for transmission and storage. High-throughput compression algorithms are therefore essential. In the sPHENIX experiment taking data at the Relativistic Heavy Ion Collider, a time projection chamber records three-dimensional (3D) particle trajectories that are highly sparse, making conventional learning-free lossy compression ineffective. Convolutional neural networks have surpassed traditional methods in compression ratio and accuracy. However, they fail to exploit sparsity for efficiency. To address these gaps, we present BCAE-VS, a bicephalous convolutional autoencoder with variable compression ratio for sparse data, which adapts compression to input complexity through key-point identification and sparse convolution. BCAE-VS achieves higher accuracy and compression ratios than prior neural approaches while being orders of magnitude smaller. Moreover, its throughput increases with sparsity—a property not observed in other methods. Although it was developed for collider experiments, BCAE-VS readily extends to other sparse data domains, such as light detection and ranging (LiDAR) sensing and 3D microscopy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Holographic Compact Disk Read-Only Memories

Compact disk read-only memories (CD-ROMs) of proposed type store digital data in volume holograms instead of in surface differentially reflective elements. Holographic CD-ROM consist largely of parts similar to those used in conventional CD-ROMs. However, achieves 10 or more times data-storage capacity and throughput by use of wavelength-multiplexing/volume-hologram scheme.

Liu, Tsuen-Hsi↗

Miniaturized Airborne Imaging Central Server System

In recent years, some remote-sensing applications require advanced airborne multi-sensor systems to provide high performance reflective and emissive spectral imaging measurement rapidly over large areas. The key or unique problem of characteristics is associated with a black box back-end system that operates a suite of cutting-edge imaging sensors to collect simultaneously the high throughput reflective and emissive spectral imaging data with precision georeference. This back-end system needs to be portable, easy-to-use, and reliable with advanced onboard processing. The innovation of the black box backend is a miniaturized airborne imaging central server system (MAICSS). MAICSS integrates a complex embedded system of systems with dedicated power and signal electronic circuits inside to serve a suite of configurable cutting-edge electro- optical (EO), long-wave infrared (LWIR), and medium-wave infrared (MWIR) cameras, a hyperspectral imaging scanner, and a GPS and inertial measurement unit (IMU) for atmospheric and surface remote sensing. Its compatible sensor packages include NASA s 1,024 1,024 pixel LWIR quantum well infrared photodetector (QWIP) imager; a 60.5 megapixel BuckEye EO camera; and a fast (e.g. 200+ scanlines/s) and wide swath-width (e.g., 1,920+ pixels) CCD/InGaAs imager-based visible/near infrared reflectance (VNIR) and shortwave infrared (SWIR) imaging spectrometer. MAICSS records continuous precision georeferenced and time-tagged multisensor throughputs to mass storage devices at a high aggregate rate, typically 60 MB/s for its LWIR/EO payload. MAICSS is a complete stand-alone imaging server instrument with an easy-to-use software package for either autonomous data collection or interactive airborne operation. Advanced multisensor data acquisition and onboard processing software features have been implemented for MAICSS. With the onboard processing for real time image development, correction, histogram-equalization, compression, georeference, and data organization, fast aerial imaging applications, including the real time LWIR image mosaic for Google Earth, have been realized for NASA fs LWIR QWIP instrument. MAICSS is a significant improvement and miniaturization of current multisensor technologies. Structurally, it has a complete modular and solid-state design. Without rotating hard drives and other moving parts, it is operational at high altitudes and survivable in high-vibration environments. It is assembled from a suite of miniaturized, precision-machined, standardized, and stackable interchangeable embedded instrument modules. These stackable modules can be bolted together with the interconnection wires inside for the maximal simplicity and portability. Multiple modules are electronically interconnected as stacked. Alternatively, these dedicated modules can be flexibly distributed to fit the space constraints of a flying vehicle. As a flexibly configurable system, MAICSS can be tailored to interface a variety of multisensor packages. For example, with a 1,024x1,024 pixel LWIR and a 8,984x6,732 pixel EO payload, the complete MAICSS volume is approximately 7x9x11 in. (=18x23x28 cm), with a weight of 25 lb (=11.4 kg).

Sun, Xiuhong↗

Simplifying Satellite and Ground Data Validation with Level-2 Subsetting

We demonstrate that scientists can simplify their satellite data validation workflow with the use of NASA Godddard Earth Sciences Data and Information Services Center (GES DISC) subsetting services. We perform a sample validation of Aura ozone products collocated with ground-based ozone measurements using subsetting services to trim satellite data to only the relevant user-defined variables and spatio-temporal region. Because the subsetting service automatically returns only relevant data granules that adhere to a set of user-defined coincidence criteria, user workload is greatly reduced. Moreover, the resultant data files are substantially smaller than full data granules due to the subsetting service further culling the data to the relevant geospatio-temporal coincidence criteria, user-defined variables, and user-defined dimensions of variables. This decreases data download throughput and file storage requirements. The validation presented here quantifies the time and file size savings that can be achieved by utilizing subsetting services within the satellite data validation workflow.

Johnson, James↗

Fast 2D Bicephalous Convolutional Autoencoder for Compressing 3D Time Projection Chamber Data

High-energy large-scale particle colliders produce data at high speed in the order of 1 terabytes per second in nuclear physics and petabytes per second in high energy physics. Developing real-time data compression algorithms to reduce such data at high throughput to fit permanent storage has drawn increasing attention. Specifically, at the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider (RHIC), a time projection chamber is used as the main tracking detector, which records particle trajectories in a volume of three-dimensional (3D) cylinder. The resulting data are usually very sparse with occupancy around 10.8%. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. The 3D convolutional neural network (CNN)-based approach, Bicephalous Convolutional Autoencoder (BCAE), outperforms traditional methods both in compression rate and reconstruction accuracy. BCAE can also utilize the computation power of graphical processing units suitable for deployment in a modern heterogeneous highperformance computing environment. This work introduces two BCAE variants: BCAE++ and BCAE-2D. BCAE++ achieves a 15% better compression ratio and a 77% better reconstruction accuracy measured in mean absolute error compared with BCAE. BCAE-2D treats the radial direction as the channel dimension of an image, resulting in a 3× speedup in compression throughput. In addition, we demonstrate an unbalanced autoencoder with a larger decoder can improve reconstruction accuracy without significantly sacrificing throughput. Lastly, we observe both the BCAE++ and BCAE-2D can benefit more from using half-precision mode in throughput (76 - 79% increase) without loss in reconstruction accuracy. The source code and links to data and pretrained models can be found at https://github.com/BNL-DAQ-LDRD/NeuralCompression_v2

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Machine Learning and IAST-Aided High-Throughput Screening of Cationic and Silica Zeolites for Alkane Capture, Storage, and Separations

We present an approach for quantitatively predicting the temperature-dependent single-component adsorption behavior of linear alkanes in silica and Na-exchanged cationic zeolites using machine learning (ML) models trained from extensive molecular simulations based on force fields with coupled cluster accuracy. A high-performing classification model was developed to distinguish between instances with negligible and non-negligible adsorption. Subsequently, two ML models were trained to predict the single-component adsorption loading and the heat of adsorption at any pressure at 300 K for any zeolite topology and silicon-to-aluminum ratio. The ML models were trained on International Zeolite Association (IZA) zeolites, and their transferability to hypothetical zeolites was successfully validated. We then expand the power of these predictions to adsorbed mixtures at arbitrary temperatures by integrating them with the Clausius–Clapeyron equation and ideal adsorbed solution theory (IAST). This approach was validated and then applied to a temperature swing adsorption separation process to demonstrate its practical utility. We demonstrate how predictions from this ML-enabled approach can allow the selection of high-performing materials that are then validated using detailed molecular simulations based on quantitatively accurate force fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sim-Situ: A Framework for the Faithful Simulation of in situ Processing

The amount of data generated by numerical simulations in various scientific domains led to a fundamental redesign of how the analysis and visualization of simulation outputs are performed. The throughput and capacity of storage subsystems have not evolved as fast as the computing power in extreme-scale supercomputers, making the classical post-hoc approach highly inefficient. In situ processing has then emerged as a solution in which simulation and data analysis/visualization are intertwined for better performance and greater interactivity.Determining the best allocation, i.e., how many resources to allocate to simulation and analysis respectively, mapping, i.e., where and at which frequency to run the analysis/visualization, and data transfer mode is a complex task whose performance assessment is crucial to the efficient execution of in situ processing. However, such a performance evaluation of different strategies usually relies either on directly running them on the targeted execution environments, which can rapidly become extremely time- and resource-consuming, or on resorting to simplified models of the components of an in situ application, which can lack of realism. In both cases, the validity of the performance evaluation is limited.In this paper, we present Sim-Situ, a simulation-based framework for the faithful performance evaluation of in situ processing strategies. We designed Sim-Situ to reflect the typical features of in situ processing systems. Thanks to its modular design, Sim-situ has the necessary flexibility to easily and faithfully evaluate the behavior and performance of various allocation, mapping, and data transfer strategies. We illustrate the simulation capabilities of Sim-Situ on a Molecular Dynamics use case. We study the impact of different strategies on performance and show how users can leverage Sim-Situ to determine interesting tradeoffs when adding analysis/visualization components to their application.

Honoré, Valentin↗

Global Change Data Center: Mission, Organization, Major Activities, and 2001 Highlights

Rapid efficient access to Earth sciences data is fundamental to the Nation's efforts to understand the effects of global environmental changes and their implications for public policy. It becomes a bigger challenge in the future when data volumes increase further and missions with constellations of satellites start to appear. Demands on data storage, data access, network throughput, processing power, and database and information management are increased by orders of magnitude, while budgets remain constant and even shrink. The Global Change Data Center's (GCDC) mission is to provide systems, data products, and information management services to maximize the availability and utility of NASA's Earth science data. The specific objectives are (1) support Earth science missions be developing and operating systems to generate, archive, and distribute data products and information; (2) develop innovative information systems for processing, archiving, accessing, visualizing, and communicating Earth science data; and (3) develop value-added products and services to promote broader utilization of NASA Earth Sciences Enterprise (ESE) data and information. The ultimate product of GCDC activities is access to data and information to support research, education, and public policy.

Wharton, Stephen W.↗

Global Change Data Center: Mission, Organization, Major Activities, and 2003 Highlights

Rapid, efficient access to Earth sciences data from satellites and ground validation stations is fundamental to the nation's efforts to understand the effects of global environmental changes and their implications for public policy. It becomes a bigger challenge in the future when data volumes increase from current levels to terabytes per day. Demands on data storage, data access, network throughput, processing power, and database and information management are increased by orders of magnitude, while budgets remain constant and even shrink.The Global Change Data Center's (GCDC) mission is to develop and operate data systems, generate science products, and provide archival and distribution services for Earth science data in support of the U.S. Global Change Program and NASA's Earth Sciences Enterprise. The ultimate product of the GCDC activities is access to data to support research, education, and public policy.

FROM↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Striped tertiary storage arrays

Data stripping is a technique for increasing the throughput and reducing the response time of large access to a storage system. In striped magnetic or optical disk arrays, a single file is striped or interleaved across several disks; in a striped tape system, files are interleaved across tape cartridges. Because a striped file can be accessed by several disk drives or tape recorders in parallel, the sustained bandwidth to the file is greater than in non-striped systems, where access to the file are restricted to a single device. It is argued that applying striping to tertiary storage systems will provide needed performance and reliability benefits. The performance benefits of striping for applications using large tertiary storage systems is discussed. It will introduce commonly available tape drives and libraries, and discuss their performance limitations, especially focusing on the long latency of tape accesses. This section will also describe an event-driven tertiary storage array simulator that is being used to understand the best ways of configuring these storage arrays. The reliability problems of magnetic tape devices are discussed, and plans for modeling the overall reliability of striped tertiary storage arrays to identify the amount of error correction required are described. Finally, work being done by other members of the Sequoia group to address latency of accesses, optimizing tertiary storage arrays that perform mostly writes, and compression is discussed.

Drapeau, Ann L.↗