Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hierarchical Data Format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop↗

Development of Secondary Archive System at Goddard Space Flight Center Version 0 Distributed Active Archive Center

The Goddard Space Flight Center (GSFC) version 0 Distributed Active Archive Center (DAAC) has been developed to support existing and pre Earth Observing System (EOS) Earth science datasets, facilitate the scientific research, and test EOS data and information system (EOSDIS) concepts. To ensure that no data is ever lost, each product received at GSFC DAAC is archived on two different media, VHS and digital linear tape (DLT). The first copy is made on VHS tape and is under the control of UniTree. The second and third copies are made to DLT and VHS media under a custom built software package named 'Archer'. While Archer provides only a subset of the functions available with commercial software like UniTree, it supports migration between near-line and off-line media and offers much greater performance and flexibility to satisfy the specific needs of a data center. Archer is specifically designed to maximize total system throughput, rather than focusing on the turn-around time for individual files. The commercial off the shelf software (COTS) hierarchical storage management (HSM) products evaluated were mainly concerned with transparent, interactive, file access to the end-user, rather than a batch-orientated, optimizable (based on known data file characteristics) data archive and retrieval system. This is critical to the distribution requirements of the GSFC DAAC where orders for 5000 or more files at a time are received. Archer has the ability to queue many thousands of file requests and to sort these requests into internal processing schedules that optimize overall throughput. Specifically, mount and dismount, tape load and unload cycles, and tape motion are minimized. This feature did not seem to be available in many COTS pacages. Archer also uses a generic tar tape format that allows tapes to be read by many different systems rather than the proprietary format found in most COTS packages. This paper discusses some of the specific requirements at GSFC DAAC, the motivations for implementing the Archer system, and presents a discussion of the Archer design that resulted.

Sherman, Mark↗

Galaxy Cruise: Deep Insights into Interacting Galaxies in the Local Universe

Abstract We present the first results from GALAXY CRUISE, a community (or citizen) science project based on data from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). The current paradigm of galaxy evolution suggests that galaxies grow hierarchically via mergers, but our observational understanding of the role of mergers is still limited. The data from HSC-SSP are ideally suited to improve our understanding with improved identifications of interacting galaxies thanks to the superb depth and image quality of HSC-SSP. We launched a community science project, GALAXY CRUISE, in 2019 and have collected over two million independent classifications of 20686 galaxies at z < 0.2. We first characterize the accuracy of the participants’ classifications and demonstrate that it surpasses previous studies based on shallower imaging data. We then investigate various aspects of interacting galaxies in detail. We show that there is a clear sign of enhanced activities of super-massive black holes and star formation in interacting galaxies compared to those in isolated galaxies. The enhancement seems particularly strong for galaxies undergoing violent mergers. We also show that the mass growth rate inferred from our results is roughly consistent with the observed evolution of the stellar mass function. The second season of GALAXY CRUISE is currently underway and we conclude with future prospects. We make the morphological classification catalog used in this paper publicly available at the GALAXY CRUISE website, which will be particularly useful for machine-learning applications.

Tanaka, Masayuki↗

HYDRA : High-speed simulation architecture for precision spacecraft formation simulation

e Hierarchical Distributed Reconfigurable Architecture- is a scalable simulation architecture that provides flexibility and ease-of-use which take advantage of modern computation and communication hardware. It also provides the ability to implement distributed - or workstation - based simulations and high-fidelity real-time simulation from a common core. Originally designed to serve as a research platform for examining fundamental challenges in formation flying simulation for future space missions, it is also finding use in other missions and applications, all of which can take advantage of the underlying Object-Oriented structure to easily produce distributed simulations. Hydra automates the process of connecting disparate simulation components (Hydra Clients) through a client server architecture that uses high-level descriptions of data associated with each client to find and forge desirable connections (Hydra Services) at run time. Services communicate through the use of Connectors, which abstract messaging to provide single-interface access to any desired communication protocol, such as from shared-memory message passing to TCP/IP to ACE and COBRA. Hydra shares many features with the HLA, although providing more flexibility in connectivity services and behavior overriding.

formation flying↗

FilDReaMS: II. Application to the analysis of the relative orientations between filaments and the magnetic field in four Herschel fields

Context. Both simulations and observations of the interstellar medium show that the study of the relative orientations between filamentary structures and the magnetic field can bring new insight into the role played by magnetic fields in the formation and evolution of filaments and in the process of star formation. Aims. We provide a first application of FilDReaMS, the new method presented in the companion paper to detect and analyze filaments in a given image. The method relies on a template that has the shape of a rectangular bar with variable width. Our goal is to investigate the relative orientations between the detected filaments and the magnetic field. Methods. We apply FilDReaMS to a small sample of four Herschel fields (G210, G300, G82, G202) characterized by different Galactic environments and different evolutionary stages. First, we look for the most prevalent bar widths, and we examine the networks formed by filaments of different bar widths as well as their hierarchical organization. Second, we compare the filament orientations to the magnetic field orientation inferred from Planck polarization data and, for the first time, we study the statistics of the relative orientation angle as functions of both spatial scale and H2 column density. Results. We find preferential relative orientations in the four Herschel fields: small filaments with low column densities tend to be slightly more parallel than perpendicular to the magnetic field; in contrast, large filaments, which all have higher column densities, are oriented nearly perpendicular (or, in the case of G202, more nearly parallel) to the magnetic field. In the two nearby fields (G210 and G300), we observe a transition from mostly parallel to mostly perpendicular relative orientations at an H 2 column density ≃ 1.1 × 10 21 cm -2 and 1.4 × 10 21 cm -2 , respectively, consistent with the results of previous studies. Conclusions. Our results confirm the existence of a coupling between magnetic fields at cloud scales and filaments at smaller scale. They also illustrate the potential of combining Herschel and Planck observations, and they call for further statistical analyses with our dedicated method.

79 ASTRONOMY AND ASTROPHYSICS↗

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]↗

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William↗

Meteorological and hydrological parameters for 17 locations of meteorological stations of the East River Watershed

The Data Package includes a set of csv files of input meteorological parameters for locations of 17 meteorological stations within the East River watershed. These meteorological datasets were downloaded from the (1) PRISM database--monthly precipitation, air temperature (minimum, mean, and maximum), vapor pressure deficit (minimum and maximum), and dewpoint temperature, and (2) NCEP/NCAR Reanalysis database – wind database. The datasets were used for calculations of the Potential Evapotranspiration (ETo), Actual Evapotranspiration (ET), Standard Precipitation Index (SPI) , Standard Evapotranspiration-Precipitation Index (SPEI) for the period from 1966 to 2021. The main research questions addressed are: the evaluation of the long-term temporal trends of climatic parameters, hierarchical clustering, and areal mapping/zonation of the East River watershed. Calculations were conducted in the Rstudio environment. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.The input datasets were downloaded from (a) PRISM database (the Northwest Alliance for Computational Science and Engineering at the Oregon State University), and (b) NCEP/NCAR Reanalysis database.

54 ENVIRONMENTAL SCIENCES↗

Hierarchical Speed Planner for Automated Vehicles: A Framework for Lagrangian Variable Speed Limit in Mixed-Autonomy Traffic

Here, this article presents a novel hierarchical speed planning framework for variable speed limits in mixed-autonomy traffic environments, leveraging server-side macroscopic control and vehicle-side microscopic execution. The framework integrates real-time traffic state estimation (TSE) and reinforcement learning (RL)-based control to mitigate congestion and improve traffic flow. A TSE enhancement module combines macroscopic data from sources like INRIX with high-resolution observations from connected autonomous vehicles (CAVs), enabling predictive modeling to address latency and noise. The target speed design module employs kernel smoothing and a buffer zone strategy to optimize traffic density and flow around bottlenecks. The proposed system was validated in the largest open-road test to date with 100 CAVs, demonstrating an overall 8% traffic density decrease, with a specific decrease of 7% upstream, 10% downstream, and a 52% decrease during the congestion formation phase at bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Covalently Linked, Two-Dimensional Quantum Dot Assemblies

Using nanoscale building blocks to construct hierarchical materials is a radical new branch point in materials discovery that promises new structures and emergent functionality. Understanding the design principles that govern nanoparticle assembly is critical to moving this field forward. By exploiting mixed ligand environments to target patchy nanoparticle surfaces, we have demonstrated a novel method of colloidal quantum dot (QD) assembly that gives rise to 2D structures. The equilibration of solutions of spherical and quasispherical QDs, including CdS, CdSe, and InP, with 2,2'-bipyridine-5,5'-diacrylic acid resulted in the preferential formation of 2D assemblies over the course of days as determined by transmission electron microscopy analysis. Small-angle X-ray scattering confirms the existence of the QD assemblies in solution. The dependence of the assembly on linker properties (length and rigidity), linker concentration, and total concentration was investigated, together with the data point to a mechanism involving ligand redistribution to create a patchy surface that maximizes the steric repulsion of neighboring QDs. By operating in an underexchanged regime, the arising patchiness results in enthalpically preferred directions of cross-linking that can be accessed by thermal equilibration.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fragmented molecular complexes: The role of the magnetic field in feeding internal supersonic motions

A hierarchical structure for molecular complexes in their cold phase i.e., preceeding the formation of massive stars, was derived from extensive large scale CO(13)(J=1=0) observations: the mass is found to be distributed into virialized clouds which fill only a very low fraction approx. 01 of the volume of the complex and are supported against gravity by internal supersonic motions. An efficient mechanism was found to transfer kinetic energy from the orbital motions of the clouds to their internal random motions. The large perturbations of the magnetic field induced at the cloud boundaries by their interactions with their neighbors generate systems of hydromagnetic waves trapped inside the clouds. The magnetic field lines being closely coupled to the gas at the densities which prevail in the bulk of the clouds volume, internal velocity dispersion is thus generated. Some conclusions derived from this data are given.

Falgarone, E.↗

Untangling the Galaxy. II. Structure within 3 kpc

We present the results of the hierarchical clustering analysis of the Gaia DR2 data to search for clusters, comoving groups, and other stellar structures. The current paper builds on the sample from the previous work, extending it in distance from 1 to 3 kpc and increasing the number of identified structures up to 8292. To aid in the analysis of the population properties, we developed a neural network called Auriga to robustly estimate the age, extinction, and distance of a stellar group based on the input photometry and parallaxes of the individual members. We apply Auriga to derive the properties of not only the structures found in this paper, but also previously identified open clusters. Through this work, we examine the temporal structure of the spiral arms. Specifically, we find that the Sagittarius Arm has moved by >500 pc in the last 100 Myr and the Perseus Arm has been experiencing a relative lull in star formation activity over the last 25 Myr. We confirm the findings of the previous paper on the transient nature of the spiral arms, with the timescale of transition of a few 100 Myr. Finally, we find a peculiar ~1 Gyr old stream of stars that appears to be heliocentric. Its origin is unclear.

79 ASTRONOMY AND ASTROPHYSICS↗

Locating Biodiversity Data Through The Global Change Master Directory

The Global Change Master Directory (GCMD) presently holds descriptions for almost 7000 data sets held worldwide. The directory's primary purpose is for data discovery. The information provided through the GCMD's Directory Interchange Format (DIF) is the set of information that a researcher would need to determine if a particular data set could be of value. By offering data set descriptions worldwide in many scientific disciplines - including meteorology, oceanography, ecology, geology, hydrology, geophysics, remote sensing, paleoclimate, solar-terrestrial physics, and human dimensions of climate change - the GCMD simplifies the discovery of data sources. Direct linkages to many of the data sets are also provided. In addition, several data set registration tools are offered for populating the directory. To search the directory, one may choose the Guided Search or Free-Text Search. Two experimental interfaces were also made available with the latest software release - one based on a keyword search and another based on a graphical interface. The graphical interface was designed in collaboration with the Human Computer Interaction Laboratory at the University of Maryland. The latest version of the software, Version 6, was released in April, 1998. It features the implementation of a scheme to handle hierarchical data set collections (parent-child relationships); a hierarchical geospatial location search scheme; a Java-based geographic map for conducting geospatial searches; a Related-URL field for project-related data set collections, metadata extensions (such as more detailed inventory information), etc.; a new implementation of the Isite software; a new dataset language field; hyperlinked email addresses, and more. The key to the continued evolution of the GCMD is in the flexibility of the GCMD database, allowing modifications and additions to made relatively easily to maintain currency, thus providing the ability to capitalize on current technology while importing all existing records. Changes are discussed and approved through an online "interoperability" forum. The next major release of the GCMD is scheduled for early 1999 and will include the incorporation of a new matrix-based interface, a rapid valids-based query system; improvement in the operations facility - important for future distributed options; new streamlined code for greater performance and maintainability; improvements in the handling of seven current fields proposed through the interoperability forum (at no expense to the data providers); and the release of DOCmorph, a more robust version of DIFmorph to translate many 'standards' multi-directionally. Issues and actions will also be addressed.

Olsen, Lola M.↗

Integrate Latimer Controls' Solution into RTAC (CRADA Final Report, CRD-23-24672)

Latimer Controls, Inc. was awarded two vouchers under the Department of Energy's American-Made Solar Prize Round 6 to conduct collaborative research at a national laboratory. The National Renewable Energy Laboratory (NREL) was selected as a partner to assist Latimer Controls in the performance evaluation of its photovoltaic (PV) control software. This collaboration focuses on developing a hardware-in-the-loop (HIL) testbed at NREL, which will be used to test and validate the Latimer PV control technology in a realistic yet de-risked environment. Both Latimer and NREL teams will work together to analyze the collected test data, derive insights, and disseminate the scientific findings. Recent studies underscore the potential of solar energy as a zero-marginal-cost and zero-emission flexibility resource within the bulk power system, particularly when integrated with advanced control systems. To enhance the performance of such systems, Latimer Controls has developed leading-edge technologies, including machine learning (ML) algorithms and hierarchical inverter set-point allocation methods. These innovations are designed to estimate the operational headroom of large PV plants for grid integration and control. However, comprehensive validation under real-world conditions remains necessary. To address this gap, the concurrent CRADA project proposes the real-world application and validation of the Latimer Control solution within a HIL environment. Initially, the Latimer algorithm was developed and tested within MATLAB Simulink, a platform suitable for research-level simulations and iterative development. However, transitioning this technology to a real solar site as an industry-ready solution necessitates implementation in a format compatible with widely used solar power plant controllers. In this additional CRADA work, the MATLAB Simulink-based logic will be translated into Structured Text, a programming language compliant with IEC 61131 standards, which is commonly used for custom logic implementations in industry-leading programmable logic controllers (PLCs), such as the Schweitzer SEL real-time automation controller (RTAC). This transition will facilitate the deployment of the Latimer Control solution in real-world solar power plants, thereby advancing the technology towards commercialization.

14 SOLAR ENERGY↗

Testing higher-order Lagrangian perturbation theory against numerical simulation. 1: Pancake models

We present results showing an improvement of the accuracy of perturbation theory as applied to cosmological structure formation for a useful range of quasi-linear scales. The Lagrangian theory of gravitational instability of an Einstein-de Sitter dust cosmogony investigated and solved up to the third order is compared with numerical simulations. In this paper we study the dynamics of pancake models as a first step. In previous work the accuracy of several analytical approximations for the modeling of large-scale structure in the mildly non-linear regime was analyzed in the same way, allowing for direct comparison of the accuracy of various approximations. In particular, the Zel'dovich approximation (hereafter ZA) as a subclass of the first-order Lagrangian perturbation solutions was found to provide an excellent approximation to the density field in the mildly non-linear regime (i.e. up to a linear r.m.s. density contrast of sigma is approximately 2). The performance of ZA in hierarchical clustering models can be greatly improved by truncating the initial power spectrum (smoothing the initial data). We here explore whether this approximation can be further improved with higher-order corrections in the displacement mapping from homogeneity. We study a single pancake model (truncated power-spectrum with power-spectrum with power-index n = -1) using cross-correlation statistics employed in previous work. We found that for all statistical methods used the higher-order corrections improve the results obtained for the first-order solution up to the stage when sigma (linear theory) is approximately 1. While this improvement can be seen for all spatial scales, later stages retain this feature only above a certain scale which is increasing with time. However, third-order is not much improvement over second-order at any stage. The total breakdown of the perturbation approach is observed at the stage, where sigma (linear theory) is approximately 2, which corresponds to the onset of hierarchical clustering. This success is found at a considerable higher non-linearity than is usual for perturbation theory. Whether a truncation of the initial power-spectrum in hierarchical models retains this improvement will be analyzed in a forthcoming work.

Buchert, T.↗

Enhancement of double-close-binary quadruples

ABSTRACT Double-close-binary quadruples (2 + 2 systems) are hierarchical systems of four stars where two short-period binary systems move around their common centre of mass on a wider orbit. Using Gaia Early Data Release 3, we search for comoving pairs where both components are eclipsing binaries. We present eight 2 + 2 quadruple systems with inner orbital periods of <0.4 d and with outer separations of ≳1000 au. All of these systems but one are newly discovered by this work, and we catalogue their orbital information measured from their light curves. We find that the occurrence rate of 2 + 2 quadruples is 7.3 ± 2.6 times higher than what is expected from random pairings of field stars. At most a factor of ∼2 enhancement may be explained by the age and metallicity dependence of the eclipsing binary fraction in the field stellar population. The remaining factor of ∼3 represents a genuine enhancement of the production of short-period binaries in wide-separation (>103 au) pairs, suggesting a close-binary formation channel that may be enhanced by the presence of wide companions.

Fezenko, Gavin B. (ORCID:0000000217169430)↗

Substructure in the stellar halo near the Sun: I. Data-driven clustering in integrals-of-motion space

Context. Merger debris is expected to populate the stellar haloes of galaxies. In the case of the Milky Way, this debris should be apparent as clumps in a space defined by the orbital integrals of motion of the stars. Aims. Our aim is to develop a data-driven and statistics-based method for finding these clumps in integrals-of-motion space for nearby halo stars and to evaluate their significance robustly. Methods. We used data from Gaia EDR3, extended with radial velocities from ground-based spectroscopic surveys, to construct a sample of halo stars within 2.5 kpc from the Sun. We applied a hierarchical clustering method that makes exhaustive use of the single linkage algorithm in three-dimensional space defined by the commonly used integrals of motion energy E, together with two components of the angular momentum, L z and L ⊥ . To evaluate the statistical significance of the clusters, we compared the density within an ellipsoidal region centred on the cluster to that of random sets with similar global dynamical properties. By selecting the signal at the location of their maximum statistical significance in the hierarchical tree, we extracted a set of significant unique clusters. By describing these clusters with ellipsoids, we estimated the proximity of a star to the cluster centre using the Mahalanobis distance. Additionally, we applied the HDBSCAN clustering algorithm in velocity space to each cluster to extract subgroups representing debris with different orbital phases. Results. Our procedure identifies 67 highly significant clusters (> 3σ), containing 12% of the sources in our halo set, and 232 subgroups or individual streams in velocity space. In total, 13.8% of the stars in our data set can be confidently associated with a significant cluster based on their Mahalanobis distance. Inspection of the hierarchical tree describing our data set reveals a complex web of relations between the significant clusters, suggesting that they can be tentatively grouped into at least six main large structures, many of which can be associated with previously identified halo substructures, and a number of independent substructures. This preliminary conclusion is further explored in a companion paper, in which we also characterise the substructures in terms of their stellar populations. Conclusions. Our method allows us to systematically detect kinematic substructures in the Galactic stellar halo with a data-driven and interpretable algorithm. The list of the clusters and the associated star catalogue are provided in two tables available at the CDS.

79 ASTRONOMY AND ASTROPHYSICS↗

Petabyte Class Storage at Jefferson Lab (CEBAF)

By 1997, the Thomas Jefferson National Accelerator Facility will collect over one Terabyte of raw information per day of Accelerator operation from three concurrently operating Experimental Halls. When post-processing is included, roughly 250 TB of raw and formatted experimental data will be generated each year. By the year 2000, a total of one Petabyte will be stored on-line. Critical to the experimental program at Jefferson Lab (JLab) is the networking and computational capability to collect, store, retrieve, and reconstruct data on this scale. The design criteria include support of a raw data stream of 10-12 MB/second from Experimental Hall B, which will operate the CEBAF (Continuous Electron Beam Accelerator Facility) Large Acceptance Spectrometer (CLAS). Keeping up with this data stream implies design strategies that provide storage guarantees during accelerator operation, minimize the number of times data is buffered allow seamless access to specific data sets for the researcher, synchronize data retrievals with the scheduling of postprocessing calculations on the data reconstruction CPU farms, as well as support the site capability to perform data reconstruction and reduction at the same overall rate at which new data is being collected. The current implementation employs state-of-the-art StorageTek Redwood tape drives and robotics library integrated with the Open Storage Manager (OSM) Hierarchical Storage Management software (Computer Associates, International), the use of Fibre Channel RAID disks dual-ported between Sun Microsystems SMP servers, and a network-based interface to a 10,000 SPECint92 data processing CPU farm. Issues of efficiency, scalability, and manageability will become critical to meet the year 2000 requirements for a Petabyte of near-line storage interfaced to over 30,000 SPECint92 of data processing power.

Chambers, Rita↗