Engineering PapersSearch

SEARCH · Engineering Papers

Results for “massive datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Understanding Structure in Line-Driven Stellar Winds Using Ultraviolet Spectropolarimetry in the Time Domain

The most massive stars are thought to lose a significant fraction of their mass in a steady wind during the main-sequence and blue supergiant phases. This in turn sets the stage for their further evolution and eventual supernova, and preconditions the surrounding medium for all following events, with consequences for ISM energization, chemical enrichment, and dust formation. Understanding these processes requires accurate observational constraints on the mass-loss rates of the most luminous stars, which can also be used to test theories of stellar wind driving. In the past, mass-loss rates have been characterized via collisional emission processes such as optical Hα and free-free radio emission, but these so-called “density squared” diagnostics require correction in the presence of widespread clumping. Recent observational and theoretical evidence points to the likelihood of a ubiquitously high level of such clumping in hot-star winds, but quantifying its effects requires a deeper understanding of the complex dynamics of radiatively driven winds and their stochastic instabilities. Furthermore, large-scale structures initiating in surface anisotropies and propagating throughout the wind can also affect wind driving and alter mass-loss diagnostics. Time series spectroscopy of high resonance-line opacity in the UV, capable of high resolution and high signal-to-noise, are required to better understand these complex dynamics, and more accurately determine mass-loss rates. The proposed Polstar mission (Scowen et al. 2022, this volume) provides the necessary resolution at the Sobolev (∼10 km s −1 ) or sound-speed (∼20 km s −1 ) scale, for over three dozen bright galactic massive stars with signal-to noise an order of magnitude above that of the celebrated MEGA campaign (Massa et al. 1995) of the International Ultraviolet Explorer (IUE), via continuous observations that track propagating structures through the winds in real time. Supporting geometric constraints are provided by the polarimetric capabilities present in all the datasets of such a mission.

Polstar – NASA MIDEX

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]

The Mercury Lander Mission Concept Study: Enabling Transformative Science from the Surface of the Inner-Most Planet

As an end-member of rockyplanet formation, Mercury holds unique clues about the original distribution of elements in the earliest stages of Solar System development,as well as how planets form and evolve in close proximity to their host starsgenerally. This Mercury Lander mission concept enables in situ surface measurements that address several fundamental science questions raised by MESSENGER’s pioneering exploration of Mercury. Such measurements are needed to understand Mercury’s unique mineralogy and geo-chemistry; to characterize the structure of the planet’s proportionally massive core’s structure; to measure the planet’s active and ancient magnetic fields at the surface; to investigate the processes that alter the surface and produce the exosphere; and to provide groundtruth for current and future remote datasets

C M Ernst

Mercury Lander: Transformative Science from the Surface of the Innermost Planet

As an end-member of terrestrial planet formation, Mercury holds unique clues about the original distribution of elements in the earliest stages of solar system development and how planets and exoplanets form and evolve in close proximity to their host stars. This Mercury Lander mission concept enables in situ surface measurements that address several fundamental science questions raised by MESSENGER’s pioneering exploration of Mercury. Such measurements are needed to understand Mercury’s unique mineralogy and geochemistry; to characterize the proportionally massive core’s structure; to measure the planet’s active and ancient magnetic fields at the surface; to investigate the processes that alter the surface and produce the exosphere; and to provide ground truth for current and future remote datasets. NASA’s Planetary Mission Concept Studies (PMCS) program awarded this study to a multidisciplinary team led by Dr. Carolyn Ernst of the Johns Hopkins Applied Physics Laboratory (APL), to evaluate the feasibility of accomplishing transformative science through a New-Frontiers-class, landed mission to Mercury in the next decade. The resulting mission concept achieves one full Mercury year (~88 Earth days) of surface operations with an ambitious, high-heritage, landed science payload, corresponding well with the New Frontiers mission framework.

Mercury lander

Moving lens effect: Simulations, forecasts, and foreground mitigation

The peculiar motion of massive objects across the line of sight imprints a dipolar temperature anisotropy pattern on the cosmic microwave background known as the moving lens effect. This effect provides a unique probe of the transverse components of the peculiar velocity field, but has not yet been detected due to its small size. We implement and validate a stacking estimator for the moving lens signal using a galaxy catalog as a tracer of massive haloes combined with reconstructed velocities from the galaxy number density field. Using simulations, we forecast detection prospects for the moving lens signal from current and upcoming microwave background and galaxy surveys. Here, we demonstrate a new foreground mitigation strategy likely sufficient for current datasets, and discuss various sources of systematic error and noise. Upcoming galaxy surveys will provide high-significance statistical detections of the moving lens effect.

Astrophysical & cosmological simulations

Inferences About the Early Moon From Gravity and Topography

Recent spacecraft missions to the Moon have significantly improved our knowledge of the lunar gravity and topography fields, and have raised some new and old questions about the early lunar history. It has frequently been assumed that the shape of the Moon today reflects an earlier equilibrium state and that the Moon has retained some internal strength. Recent analysis indicating a superisostatic state of some lunar basins lends support to this hypothesis. On its simplest level, the present shape of the Moon is slightly flattened by 2.2 +/- 0.2 km while its gravity field, represented by an equipotential surface, is flattened only about 0.5 km. The hydrostatic component to the flattening arising from the Moon's present day rotation contributes only 7 m. This difference between the topographic shape of the MOon and the shape of its gravitational equipotential has frequently been explained as the "memory" of an earlier moon that was rotating faster and had a correspondingly larger hydrostatic flattening. To obtain this amount of hydrostatic flattening from rotation alone, and accounting for the contribution of the present-day gravity field, the Moon's rotation rate would need to be about 15x greater than at present, leading ot a period of < 2 days. Maintaining its synchronous rotation with Earth would require a radius for the Moon's orbit of approximately 9 Earth Radii. Unfortunately, our confidence in the observed lunar flattening is not as great as we would like. The uncertainty of .02 km may not properly reflect the limitations of the Clementine dataset, which did not sample poleward of latitudes 81 N and 79 S. Also, the large variation of topography +/- 8 km seen on the MOon dwarfs our estimate fo the flattening. Further the lunar south pole is on the edge of, or possibly inside the massive deep, South Pole-Aitken Basin. Thus, polar radii could be underestimated. This would yield a smaller flattening, which would imply a greater lunar rotation period and orbital radius. However, Basin compensation states and analyses of support and relaxation of topography at long wavelengths point to a lunar shape that has retained a flattening from an earlier faster rotation period.

Smith, D. E.

UIT support observations archive

We are in the process of archiving the approximately 1.2Tbytes of imaging data acquired in support of UIT observations. The UIT is one of three telescopes comprising the ASTRO spacecraft and is a 38-cm f/9 Ritchey-Chretien telescope with wavelength coverage 1200 A to 3000 A and a 40 arcminute diameter field-of-view at 4 min FWHM resolution. During the ASTRO shuttle mission, there were 65 different pointings (some with multiple targets) of the UIT and hence 65 fields. Our support data of these fields were obtained with the KPNO and CTIO 0.9m telescopes and several of the 2048x2048 CCD's. Our images are typically about 20 acrminutes on a side and contain the entire UV target except for nearby galaxies, for which mosaiced images were created. All good quality images in the broadband and narrow band filters for every target are being archived. Though all of the UIT targets were well studied astronomical objects, these frames are many of the first large field format images of them and, when combined with the soon to be released UV frames, provide a unique dataset. These data were already used to address a wide variety of astronomical questions. Vacuum ultraviolet (VUV) observations were used to study star formation in a sample of nearby galaxies, since integrated VUV - optical colors provide the most sensitive available measure of the formation rate of massive stars. The blue stages later in stellar evolution are also being studied. These data allow more accurate determination of the helium and metallicities of hothorzontal branch stars. These data are being prepared to release to the public in CD-ROM format through the National Space Science Data Center.

Smith, Eric P.

Value added data archiving

Researchers in the Molecular Sciences Research Center (MSRC) of Pacific Northwest Laboratory (PNL) currently generate massive amounts of scientific data. The amount of data that will need to be managed by the turn of the century is expected to increase significantly. Automated tools that support the management, maintenance, and sharing of this data are minimal. Researchers typically manage their own data by physically moving datasets to and from long term storage devices and recording a dataset's historical information in a laboratory notebook. Even though it is not the most efficient use of resources, researchers have tolerated the process. The solution to this problem will evolve over the next three years in three phases. PNL plans to add sophistication to existing multilevel file system (MLFS) software by integrating it with an object database management system (ODBMS). The first phase in the evolution is currently underway. A prototype system of limited scale is being used to gather information that will feed into the next two phases. This paper describes the prototype system, identifies the successes and problems/complications experienced to date, and outlines PNL's long term goals and objectives in providing a permanent solution.

Berard, Peter R.

Harnessing on-machine metrology data for prints with a surrogate model for laser powder directed energy deposition

In this study, we leverage the massive amount of multi-modal on-machine metrology data generated from Laser Powder Directed Energy Deposition (LP-DED) to construct a comprehensive surrogate model of the 3D printing process. By employing Dynamic Mode Decomposition with Control (DMDc), a data-driven technique, we capture the complex physics inherent in this extensive dataset. This physics-based surrogate model emphasizes thermodynamically significant quantities, enabling us to accurately predict key process outcomes. The model ingests 21 process parameters, including laser power, scan rate, and position, while providing outputs such as melt pool temperature, melt pool size, and other essential observables. Furthermore, it incorporates uncertainty quantification to provide bounds on these predictions, enhancing reliability and confidence in the results. We then deploy the surrogate model on a new, unseen part and monitor the printing process as validation of the method. Our experimental results demonstrate that the predictions align with actual measurements with high accuracy, confirming the effectiveness of our approach. Furthermore, this methodology not only facilitates real-time predictions but also operates at process-relevant speeds, establishing a basis for implementing feedback control in LP-DED.

Digital twins

The stellar origins of 96 Zr excesses in presolar graphites from the Murchison meteorite

Context. Zirconium-96 is a stable isotope that can be synthesized under different neutron-rich nucleosynthetic conditions. Astrophysical models predict its production to occur in various stellar environments: from low-to-intermediate-mass asymptotic giant branch (AGB) stars to massive stars and core-collapse supernovae. Aims. Detections of 96Zr excesses, in combination with other isotopic measurements from presolar grains can provide unique constraints on its stellar origin. Presolar grains are microscopic particles found in primitive Solar System materials, which formed in stellar winds and supernova ejecta. The isotopic composition of each grain can provide us a snapshot of the nucleosynthetic processes that took place during the parent star’s lifetime. Methods. In this study, we measured the stable isotopes of C, N, O, Mo, Zr, and Ru in high-density presolar graphite grains from the Murchison meteorite and found four grains that contain positive isotopic anomalies in 96 Zr carried by their internal subgrains. We analyzed multi-element isotopic datasets from each grain to explore the source of the observed 96 Zr excesses. Results. Comparisons with stellar models indicate that two grains likely condensed in an intermediate-mass AGB star with initial metallicity of Z ≤ 0.014. Their 96 Zr/ 94 Zr ratios also match those predicted for born-again AGB stars undergoing a very late thermal pulse and rapidly accreting white dwarfs. After comparing the relative populations of the aforementioned dust-producing stars, we propose rapidly accreting white dwarfs as a new, and more likely, stellar source for one of the presolar grains. The remaining two grains could have originated in the supernova ejecta of massive stars, due to correlated excesses in the p-nuclides, 92, 94 Mo. Thus, grains with 96 Zr anomalies can have a variety of stellar origins, in agreement with theoretical studies. Conclusions. Our study highlights the importance of multi-element analysis in constraining the types of stars where presolar grains have condensed. These data will help improve our understanding of various nucleosynthesis processes in different stellar phases.

79 ASTRONOMY AND ASTROPHYSICS

SPT clusters with DES and HST weak lensing. II. Cosmological constraints from the abundance of massive halos

We present cosmological constraints from the abundance of galaxy clusters selected via the thermal Sunyaev-Zel’dovich (SZ) effect in South Pole Telescope (SPT) data with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). The cluster sample is constructed from the combined SPT-SZ, SPTpol ECS, and SPTpol 500d surveys, and comprises 1,005 confirmed clusters in the redshift range 0.25–1.78 over a total sky area of 5200 deg 2 . We use DES Year 3 weak-lensing data for 688 clusters with redshifts 𝑧 < 0.95 and HST weak-lensing data for 39 clusters with 0.6 < 𝑧 < 1.7. The weak-lensing measurements enable robust mass measurements of sample clusters and allow us to empirically constrain the SZ observable-mass relation without having to make strong assumptions about, e.g., the hydrodynamical state of the clusters. For a flat Λ⁢ CDM cosmology, and marginalizing over the sum of massive neutrinos, we measure Ω m = 0.286 ± 0.032, 𝜎 8 = 0.817 ± 0.026, and the parameter combination 𝜎 8 ⁢(Ω m /0.3) 0.25 = 0.805 ± 0.016. Our measurement of 𝑆 8 ≡ 𝜎 8 ⁢$\sqrt{Ω_{m}/0.3}$ = 0.795 ± 0.029 and the constraint from Planck CMB anisotropies (2018 TT, TE, EE+lowE) differ by 1.1⁢𝜎. In combination with that Planck dataset, we place a 95% upper limit on the sum of neutrino masses ∑𝑚 𝜈 < 0.18 eV. When additionally allowing the dark energy equation of state parameter 𝑤 to vary, we obtain 𝑤 = −1.45 ± 0.31 from our cluster-based analysis. In combination with Planck data, we measure 𝑤 =−1.3⁢4$^{+0.22}_{−0.15}$, or a 2.2⁢𝜎 difference with a cosmological constant. We use the cluster abundance to measure 𝜎8 in five redshift bins between 0.25 and 1.8, and we find the results to be consistent with structure growth as predicted by the Λ⁢ CDM model fit to Planck primary CMB data.

79 ASTRONOMY AND ASTROPHYSICS

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)

Massively parallel reporter assays and mouse transgenic assays provide correlated and complementary information about neuronal enhancer activity

High-throughput massively parallel reporter assays (MPRAs) and phenotype-rich in vivo transgenic mouse assays are two potentially complementary ways to study the impact of noncoding variants associated with psychiatric diseases. Here, we investigate the utility of combining these assays. Specifically, we carry out an MPRA in induced human neurons on over 50,000 sequences derived from fetal neuronal ATAC-seq datasets and enhancers validated in mouse assays. We also test the impact of over 20,000 variants, including synthetic mutations and 167 common variants associated with psychiatric disorders. We find a strong and specific correlation between MPRA and mouse neuronal enhancer activity. Four out of five tested variants with significant MPRA effects affected neuronal enhancer activity in mouse embryos. Mouse assays also reveal pleiotropic variant effects that could not be observed in MPRA. Our work provides a catalog of functional neuronal enhancers and variant effects and highlights the effectiveness of combining MPRAs and mouse transgenic assays.

Kosicki, Michael

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen

Carbonate Platform Development and Stromatolite Morphogenesis: Constraints on Environmental and Biological Evolution

Work has been completed on the digital mapping of a terminal Proterozoic reef complex in Namibia. This complex formed an isolated carbonate platform developed downdip on a carbonate ramp of the Nama Group. The stratigraphic evolution of the platform was digitally reconstructed from an extensive dataset that was compiled by using digital surveying technologies. The platform comprises three accommodation cycles in which each subsequent cycle experienced progressively greater influence of a long-term accommodation increase. Aggradation and progradation during the first cycle resulted in a flat, uniform, sheet-like platform. The coarsening and shallowing-upward sequence representing the first cycle is dominated by columnar stromatolitic thrombolites and massive dolostones with interbedded mudstone-grainstone at the base of the sequence grading into cross-bedded dolostones. The second cycle features aggradation, formation of a distinct margin containing thrombolite mounds and domes, and the development of a bucket geometry. Columnar stromatolitic thrombolites dominate the platform interior. The final stage of platform development shows a deepening trend with initial aggradation and formation of well-bedded, thin deposits in the interior and mound development at the margins. While the interior drowned, the platform margin kept up with rising sea level and a complex pinnacle reef formed containing fused and coalesced thrombolite mounds flanked by bioclastic grainstones (containing Cloudina and Namacalathus fossils) and collapse breccias. A set of isolated large thrombolite mounds flanked by shales indicate the final stage of the carbonate platform. During a progressive increase in accommodation, a flat-topped isolated carbonate platform becomes aerially less extensive by either backstepping or formation of smaller pinnacles or a combination of both. The overall geometric evolution of the studied platform from flat-topped to bucket with elevated margins is recorded in many Proterozoic and Phanerozoic isolated carbonate platforms with similar dimensions. The terminal Proterozoic, microbial-dominated, isolated carbonate platform of this study clearly illustrates that the answer to accommodation changes was already familiar among carbonate platforms before the dawn of metazoan-dominated platforms.

Grotzinger, John P.

Fast Multivariate Search on Large Aviation Datasets

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which can contain up to several gigabytes of data. Surprisingly, research on MTS search is very limited. Most existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two provably correct algorithms to solve this problem (1) an R-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences, and (2) a List Based Search (LBS) algorithm which uses sorted lists for indexing. We demonstrate the performance of these algorithms using two large MTS databases from the aviation domain, each containing several millions of observations Both these tests show that our algorithms have very high prune rates (>95%) thus needing actual

Bhaduri, Kanishka

Chromosome-scale Genome Assembly of the Most Abundant Ectomycorrhizal Fungus Cenococcum Geophilum Reveals Massive TE Expansion and RIP Defense Mechanism

Transposable elements (TEs) play crucial roles in genome evolution and ecological adaptation in fungi, yet their dynamics in ectomycorrhizal species remain poorly understood. Cenococcum geophilum, the most widespread ectomycorrhizal fungus in boreal and temperate forests with its large, repeat-rich genome, represents an ideal system to investigate TE-mediated adaptation to the physical environment and symbiotic lifestyle. However, previous studies have been limited by fragmented genome assemblies that prevented the resolution of repeat-rich regions. We assembled a telomere-to-telomere reference genome of C. geophilum strain 1.58 using PacBio HiFi and Hi-C datasets, resulting in a 178.54 Mbp genome with seven contiguous chromosomes. We identified 14,145 genes and over 78% of the genome consists of transposable elements (TEs). Of these, 94% are affected by repeat-induced point mutations (RIP), a genome defense mechanism that acts during the sexual reproduction phase, indicating cryptic or ancient sexual reproduction in this putatively asexual fungus. Long terminal repeat retrotransposons, LINEs, and DNA transposons dominate, with three TE families (Ty3, Ty1, and Tad1) contributing over 60% of the genome size, indicating recent transposition bursts. Screening of 15 additional C. geophilum strains revealed recent and lineage-specific TE expansions, implying that several TEs escaped the RIP machinery and retained potential activity. Supporting TE activity in the context of symbiosis, we found 56 TEs differentially transcribed between ectomycorrhizal and free-living mycelium tissues. An even higher number (n = 66) of TEs were differentially expressed between stress resistance morphology (i.e. sclerotia) and free-living mycelium. This supports that TEs are differentially regulated as a response to symbiotic and stress-related conditions. Our results demonstrate that the C. geophilum genome expansion was driven by a few lineage-specific TE families in recent history, with high RIP activity attesting to sexual reproduction. We also provide insights how TEs could respond to lifestyle transitions and traits associated with desiccation resistance.

Cenococcum geophilum

How Well Does NASA GEOS Model Perform in Simulating Dust Deposition into the Tropical Atlantic Ocean?

Massive dust emitted from North Africa can transport long distances across the tropical Atlantic Ocean, reaching the Americas. Dust deposition along the transit adds microorganisms and essential nutrients to marine ecosystem, which has important implications for biogeochemical cycle and climate. However, assessing the dust-ecosystemclimate interactions has been hindered in part by the paucity of dust deposition measurements and large uncertainties associated with oversimplified representations of dust processes in current models. We have recently produced a unique dataset of seasonal dust deposition flux and dust loss frequency into the tropical Atlantic Ocean at a nominal resolution of 200 km x 500 km by using the decade-long (2007-2016) record of aerosol three-dimensional distribution from four satellite sensors, namely CALIOP, MODIS, MISR, and IASI. On the basis of the ten-year average, the yearly dust deposition into the tropical Atlantic Ocean is estimated at 98-153 Tg. The dust deposition shows large spatial and temporal (on seasonal and interannual scale) variability. The satellite observations also yield an estimate of annual mean dust loss frequency of 0.052 ~ 0.078 d-1, a useful diagnostic that makes it possible to disentangle the dust transport and removal processes from the dust emissions when identifying the major factors contributing to the uncertainties and biases in the model simulated dust deposition. In this study, we use the dataset along with in situ and remote sensing observations to assess how well NASA GEOS model performs in simulating trans-Atlantic dust transport and deposition. We found that the GEOS modeling of dust deposition falls within the range of satellite-based estimates. However, this reasonable agreement in dust deposition is a compensation of the model's underestimate of dust emissions and overestimate of dust removal efficiency. Further, the overestimate of dust removal efficiency results largely from the model's overestimate of rainfall rate. Our results provide insights into the model's deficiencies at process level, which could better guide model improvements.

Yu, Hongbin