Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “file transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Network issues for large mass storage requirements

File Servers and Supercomputing environments need high performance networks to balance the I/O requirements seen in today's demanding computing scenarios. UltraNet is one solution which permits both high aggregate transfer rates and high task-to-task transfer rates as demonstrated in actual tests. UltraNet provides this capability as both a Server-to-Server and Server-to-Client access network giving the supercomputing center the following advantages highest performance Transport Level connections (to 40 MBytes/sec effective rates); matches the throughput of the emerging high performance disk technologies, such as RAID, parallel head transfer devices and software striping; supports standard network and file system applications using SOCKET's based application program interface such as FTP, rcp, rdump, etc.; supports access to the Network File System (NFS) and LARGE aggregate bandwidth for large NFS usage; provides access to a distributed, hierarchical data server capability using DISCOS UniTree product; supports file server solutions available from multiple vendors, including Cray, Convex, Alliant, FPS, IBM, and others.

Perdue, James↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

Wave Tank Characterization Data

This data set was collected from a 16k gallon single paddle type wave tank after adjustments were made to the transfer function in the control software. The file names describe the commanded wave height and period which can then be compared with the actual measured height and period of the generated waves.

16 TIDAL AND WAVE POWER↗

Viability of S3 Object Storage for the ASC Program at Sandia

Recent efforts at Sandia such as DataSEA are creating search engines that enable analysts to query the institution’s massive archive of simulation and experiment data. The benefit of this work is that analysts will be able to retrieve all historical information about a system component that the institution has amassed over the years and make better-informed decisions in current work. As DataSEA gains momentum, it faces multiple technical challenges relating to capacity storage. From a raw capacity perspective, data producers will rapidly overwhelm the system with massive amounts of data. From an accessibility perspective, analysts will expect to be able to retrieve any portion of the bulk data, from any system on the enterprise network. Sandia’s Institutional Computing is mitigating storage problems at the enterprise level by procuring new capacity storage systems that can be accessed from anywhere on the enterprise network. These systems use the simple storage service, or S3, API for data transfers. While S3 uses objects instead of files, users can access it from their desktops or Sandia’s high-performance computing (HPC) platforms. S3 is particularly well suited for bulk storage in DataSEA, as datasets can be decomposed into object that can be referenced and retrieved individually, as needed by an analyst. In this report we describe our experiences working with S3 storage and provide information about how developers can leverage Sandia’s current systems. We present performance results from two sets of experiments. First, we measure S3 throughput when exchanging data between four different HPC platforms and two different enterprise S3 storage systems on the Sandia Restricted Network (SRN). Second, we measure the performance of S3 when communicating with a custom-built Ceph storage system that was constructed from HPC components. Overall, while S3 storage is significantly slower than traditional HPC storage, it provides significant accessibility benefits that will be valuable for archiving and exploiting historical data. There are multiple opportunities that arise from this work, including enhancing DataSEA to leverage S3 for bulk storage and adding native S3 support to Sandia’s IOSS library.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Transferable Output ASCII Data (TOAD) editor version 1.0 user's guide

The Transferable Output ASCII Data (TOAD) editor is an interactive software tool for manipulating the contents of TOAD files. The TOAD editor is specifically designed to work with tabular data. Selected subsets of data may be displayed to the user's screen, sorted, exchanged, duplicated, removed, replaced, inserted, or transferred to and from external files. It also offers a number of useful features including on-line help, macros, a command history, an 'undo' option, variables, and a full compliment of mathematical functions and conversion factors. Written in ANSI FORTRAN 77 and completely self-contained, the TOAD editor is very portable and has already been installed on SUN, SGI/IRIS, and CONVEX hosts.

Bingel, Bradford D.↗

Performance analysis of local area networks

A simulation of the TCP/IP protocol running on a CSMA/CD data link layer was described. The simulation was implemented using the simula language, and object oriented discrete event language. It allows the user to set the number of stations at run time, as well as some station parameters. Those parameters are the interrupt time and the dma transfer rate for each station. In addition, the user may configure the network at run time with stations of differing characteristics. Two types are available, and the parameters of both types are read from input files at run time. The parameters include the dma transfer rate, interrupt time, data rate, average message size, maximum frame size and the average interarrival time of messages per station. The information collected for the network is the throughput and the mean delay per packet. For each station, the number of messages attempted as well as the number of messages successfully transmitted is collected in addition to the throughput and mean packet delay per station.

Alkhatib, Hasan S.↗

LLNL Kimberlina 1.2 NUFT Simulations June 2018 (v2)

This dataset contains the output 6,000, 3-dimensional reactive multi-phase flow and transport aquifer simulations of brine and CO2 leakage into a protective aquiver in California’s San Joaquin Valley and input data files detailing the geologic mesh, aquifer physical properties and CO2 and brine injection rates. This data set was generated as an ongoing effort with the US DOE National Risk Assessment Partnership (NRAP) to evaluate the effectiveness of monitoring techniques to detect brine and CO2 leakage from legacy wells into underground sources of drinking water overlaying a CO2 storage reservoir. Each simulation contains a unique set of input parameters, generated stochastically. The outputs consist of these upper three geologic layers (from top): the Etchegoin, Macoma-Chanac, Santa Margarita-McLure formations. These simulations span the several distances (1, 3 and 6 km or wells W31-0.2, W31-0.5 and W31-1.0, respectively) from the CO2 injector, initiated from bottom hole pressure and saturation to calculate wellbore leakage from the storage reservoir, with low and high regional groundwater gradients and wellbore leakage into 5 leaky nodes. The dataset includes 1,000 unique simulations for each distance, which each contain a unique aquifer heterogeneity, aquifer and caprock permeability, and two model generations are included with a high permeability (prod07) and hybrid permeability (prod09). The range of permeability distributions is listed in Table 1. Each model generation consists of 3,000 simulations. Included in the dataset are the leakage rates determined from 2D wellbore models which utilize the pressure and CO2 saturation from LBL's reservoir simulations, NUFT mesh files with distributed lithology, NUFT rocktab files which describe the material properties for the geologic layers and the NUFT input files and post-processed output 'ntab' files. Each ntab file contains spatial (rows) and temporal (columns) model output tables for each model cell, the locations (x,y,z) and dimensions for each cells (dx, dy, dz). Table 1. Permeability distribution ranges for prod07 and prod09 model generations Geologic Layer: Permeability Range (log10 m^2) prod07 prod09 Etchegoin -12.92 to -10.92 -13.70 to -11.44 Macoma-Chanac -12.72 to -10.72 -13.50 to -11.24 Santa Margarita-McLure -12.70 to -10.70 -13.48 to -11.22 The input files used to generate the model include which are included in the dataset are: Time series of CO2 leakage input into the model (ex: Q_brn.W31-0.2.sim1000.layers123.tab) Time series of CO2 leakage input into the model (ex: Q_CO2.W31-0.2.sim1000.layers123.tab) Physical properties of the aquifer materials detailing the aquifer porosity, solid density, partitioning coefficients, permeabilities and van-Genuchten parameters detailed in a NUFT rocktab file: (ex: sim1000.usnt.rocktab) Numerical mesh and geologic data assigned to each model cell detailed in a NUFT genmsh format (ex: sim1000.mesh_k16.prod07.trans.genmsh) The primary output parameters are: pH (use absolute value) Change in TDS (mg/kg) Change in Pressure (Pa) Change CO2 gas saturation (fraction range 0.0-1.0) for example, the directory /p/lscratchh/mansoor1/nrap/kimberlina/prod09/mainfiles/sim1000/W31- 0.2 contains: sim1000.W31-0.2.trans.pH.red.ntab sim1000.W31-0.2.no_bg.trans.TDS.red.ntab sim1000.W31-0.2.usnt.P.deltabg.red.ntab sim1000.W31-0.2.usnt.CO2_sat.deltabg.red.ntab Each row in the NTAB files consist of model output per numerical grid cell. Each output file contains 33 columns (variables), including the information of numerical records, geologic location and sizes and the simulated parameter values over time. The first 13 variables are about numerical records and relative geologic information for a simulation grid: 1. index: simulation index 2. i: the ith grid of x-axis 3. j: the ith grid of y-axis 4. k: the ith grid of z-axis 5. element_ref: element reference 6. nuft_ind: nuft index 7. x: grid location in the x axis direction 8. y: grid location in the y axis direction 9. z: grid location in the z axis direction 10. dx: grid length in the x axis direction 11. dy: grid length in the y axis direction 12. dz: grid length in the z axis direction 13. volume: volume of the simulation grid The remainder (14, 15, 16...) variables are the simulated parameter values over time, take Pressure as an example, are: 14. 0.0y: initial pressure per cell. 15. 10.0y: simulated pressure at the end of the 10th year. 16. 20.0y: simulated pressure at the end of the 20th year. ... (repeated for every 10 years until 200 years)... The model extends 10,000 m, 5,000 m and 1,411 m in the x,y and z dimensions, respectively. The mesh consists of 164,832 cells with mesh dimensions of 101 x 51 x 32 (nx, ny, nz), with cell dimensions ranging from 100 m laterally (along x and y-axis) and model layers are as designated in the z-axis: Layer 1: atmosphere (1e-30 m thick) Layer 2: upper caprock (10 m thick) Layers 3-13: Etchegoin (536.23 m thck) Layers 14-27: Macoma-Chanac (679.04 m thick) Layers 28-32: Santa Margarita-McLure (185.94 m thick) The wellbore is placed along node i=51, j=26, and extends vertically along 5 nodes from the top to the bottom of the model. Special instructions when extracting files: Each Gzip archive (ex: prod07.sim1000-sim00099.tar.gz) contains 100 simulations. Gzip archives should be transferred into base directories (ie. In Linux: mkdir prod07; mv prod07.*.tar.gz prod07/.) before extracting, or files will be overwritten. Each sub-simulation tree should have the following file structure pattern (using the linux 'tree' command): |-- prod07 | |-- sim0001 | |-- W31-0.2 | | |-- Q_brn.W31-0.2.sim0001.layers123.tab | | |-- Q_co2.W31-0.2.sim0001.layers123.tab | | |-- sim0001.W31-0.2.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-0.2.trans.pH.red.ntab | | |-- sim0001.W31-0.2.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-0.2.usnt.P.deltabg.red.ntab | |-- W31-0.5 | | |-- Q_brn.W31-0.5.sim0001.layers123.tab | | |-- Q_co2.W31-0.5.sim0001.layers123.tab | | |-- sim0001.W31-0.5.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-0.5.trans.pH.red.ntab | | |-- sim0001.W31-0.5.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-0.5.usnt.P.deltabg.red.ntab | |-- W31-1.0 | | |-- Q_brn.W31-1.0.sim0001.layers123.tab | | |-- Q_co2.W31-1.0.sim0001.layers123.tab | | |-- sim0001.W31-1.0.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-1.0.trans.pH.red.ntab | | |-- sim0001.W31-1.0.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-1.0.usnt.P.deltabg.red.ntab | |-- sim0001.mesh_k16.prod07.trans.genmsh Disclaimer This document was prepared as an account of work sponsored by an agency of the United States government. Neither the United States government nor Lawrence Livermore National Security, LLC, nor any of their employees makes any warranty, expressed or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States government or Lawrence Livermore National Security, LLC. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States government or Lawrence Livermore National Security, LLC, and shall not be used for advertising or product endorsement purposes. Lawrence Livermore National Laboratory is operated by Lawrence Livermore National Security, LLC, for the U.S. Department of Energy, National Nuclear Security Administration under Contract DE-AC52-07NA27344. This report was reviewed and released as LLNL-MI-753464.

aquifer↗

Characterizing parallel file-access patterns on a large-scale multiprocessor

Rapid increases in the computational speeds of multiprocessors have not been matched by corresponding performance enhancements in the I/O subsystem. To satisfy the large and growing I/O requirements of some parallel scientific applications, we need parallel file systems that can provide high-bandwidth and high-volume data transfer between the I/O subsystem and thousands of processors. Design of such high-performance parallel file systems depends on a thorough grasp of the expected workload. So far there have been no comprehensive usage studies of multiprocessor file systems. Our CHARISMA project intends to fill this void. The first results from our study involve an iPSC/860 at NASA Ames. This paper presents results from a different platform, the CM-5 at the National Center for Supercomputing Applications. The CHARISMA studies are unique because we collect information about every individual read and write request and about the entire mix of applications running on the machines. The results of our trace analysis lead to recommendations for parallel file system design. First the file system should support efficient concurrent access to many files, and I/O requests from many jobs under varying load conditions. Second, it must efficiently manage large files kept open for long periods. Third, it should expect to see small requests predominantly sequential access patterns, application-wide synchronous access, no concurrent file-sharing between jobs appreciable byte and block sharing between processes within jobs, and strong interprocess locality. Finally, the trace data suggest that node-level write caches and collective I/O request interfaces may be useful in certain environments.

Purakayastha, Apratim↗

An optimal user-interface for EPIMS database conversions and SSQ 25002 EEE parts screening

The Electrical, Electronic, and Electromechanical (EEE) Parts Information Management System (EPIMS) database was selected by the International Space Station Parts Control Board for providing parts information to NASA managers and contractors. Parts data is transferred to the EPIMS database by converting parts list data to the EP1MS Data Exchange File Format. In general, parts list information received from contractors and suppliers does not convert directly into the EPIMS Data Exchange File Format. Often parts lists use different variable and record field assignments. Many of the EPES variables are not defined in the parts lists received. The objective of this work was to develop an automated system for translating parts lists into the EPIMS Data Exchange File Format for upload into the EPIMS database. Once EEE parts information has been transferred to the EPIMS database it is necessary to screen parts data in accordance with the provisions of the SSQ 25002 Supplemental List of Qualified Electrical, Electronic, and Electromechanical Parts, Manufacturers, and Laboratories (QEPM&L). The SSQ 2S002 standards are used to identify parts which satisfy the requirements for spacecraft applications. An additional objective for this work was to develop an automated system which would screen EEE parts information against the SSQ 2S002 to inform managers of the qualification status of parts used in spacecraft applications. The EPIMS Database Conversion and SSQ 25002 User Interfaces are designed to interface through the World-Wide-Web(WWW)/Internet to provide accessibility by NASA managers and contractors.

Watson, John C.↗

NOvA Inclusive Nue CC Cross Section Data Release

NOvA inclusive electron neutrino cross section results presented at Neutrino2020Exposure:8.09E20 protons-on-target, neutrino-enhanced beamData release contains 3 ROOT files containing the double-differential measurement with respect to electron angle (cos theta) and electron energy, a single-differential measurement with respect to squared four-momentum transfer (Q2), and the total cross section as a function of neutrino energy. The 1D cross-section measurements include the same kinematic phase space restrictions present in the double-differential analysis.The files, xsec_{dThdE/Enu/dQ2}.root, contain the extracted cross section and associated covariance matrices for the double-differential cross section with respect to the observed electron kinematics, dThdE, neutrino energy, Enu, and squared four-momentum transfer, Q^2.In each file there will be the following histograms:— TH2D/TH1D xsec: unfolded measured cross section— TH2D xsec_cov: cross section covariance matrix— TH2D stat_cov: statistical covariance matrixIncludes text files, xsec_{dThdE/Enu/dQ2}.txt show the corresponding histograms in plain text format.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multidimensional Data Aggregation in the Cloud with Application to Geostationary Satellite-based Air Quality Monitoring

Scientists use satellite data for studying Earth's systems, and the remote sensing data that these satellites collect are typically separated into files of a size small enough for efficient network transfer and storage. However, researchers usually prefer to analyze the data based on real-world dimensions like time, space, or elevation. To help with this, NASA's Atmospheric Science Data Center (ASDC) developed a new cloud-based tool that combines these smaller data chunks into larger, more useful datasets. The tool works on Network Common Data Form (netCDF4) and some HDF5 formatted files, and it is available as a service in NASA's Earthdata Cloud. In this presentation, we showcase this service using data from the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. By combining TEMPO's continuous observations over time, we create longer and more informative analysis-ready time series to facilitate the study of air quality patterns. Insights gained will provide a more comprehensive understanding of pollution sources, transport patterns, and their effects on the environment and human health.

Daniel Kaufman↗

Solar Astronomy Data Base: Packaged Information on Diskette

In its role as a library, the National Geophysical Data Center has transferred to diskette a collection of small, digital files of routinely measured solar indices for use on an IBM-compatible desktop computer. Recording these observations on diskette allows the distribution of specialized information to researchers with a wide range of expertise in computer science and solar astronomy. Every data set was made self-contained by including formats, extraction utilities, and plain-language descriptive text. Moreover, for several archives, two versions of the observations are provided - one suitable for display, the other for analysis with popular software packages. Since the files contain no control characters, each one can be modified with any text editor.

Mckinnon, John A.↗

Distributed Computing for the Project 8 Experiment

The Project 8 collaboration aims to measure the absolute neutrino mass or improve on the current limit by measuring the tritium beta decay electron spectrum. We present the current distributed computing model for the Project 8 experiment. Project 8 is in its second phase of data taking with a near continuous data rate of 1Gbps. The current computing model uses DIRAC (Distributed Infrastructure with Remote Agent Control) for its workflow and data management. A detailed meta-data assignment using the DIRAC File Catalog is used to automate raw data transfers and subsequent stages of data processing. The DIRAC system is deployed on containers managed using a Kubernetes cluster to provide a scalable infrastructure. A modified DIRAC Site Director provides the ability to submit jobs using Singularity on opportunistic High-Performance Computing (HPC) sites.

Distributed Computing, Kubernetes, DIRAC, Project ↗

Exit Presentation - Jared Ruzicka

The exit presentation provides an in depth examination of Spring 2020 NIFS intern, Jared Ruzicka’s, work on POST2 including creation of a module containing heritage aerodatabases and manual automation. The aerodatabase module incorporates a variety of legacy fortran and .dat aerodatabases into a POST2 module with example inputdecks verified by MATLAB mex files for 3 and 6 DOF simulations in nominal and dispersed conditions. The manual automation project discusses the transfer of the POST2 User’s Manual from word documents to text-based markdown files and the process through which a python script converts the manual to a PDF with improved formatting and compliance potential in a fraction of current manual generation time.

Jared Ruzicka↗