Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Synergistic Retrievals of Ice in High Clouds From Elastic Backscatter Lidar, Ku-band Radar and Submillimeter Wave Radiometer Observations

In this study, we investigate the synergy of elastic backscatter lidar, Ku-band radar, and sub-millimeter-wave radiometer measurements in the retrieval of ice from satellite observations. The synergy is analyzed through the generation of a large dataset of IceWater Content (IWC) profiles and simulated lidar, radar and radiometer observations. The characteristics of the instruments e.g. frequencies, sensitivities, etc. are set based on the expected characteristics of instruments of the Atmosphere Observing System (AOS) mission. A hold-out validation methodology is used to assess the accuracy of the IWC profiles retrieved from various combinations of observations from the three instruments. Specifically, the IWC and associated observations are randomly divided into two datasets, one for training and the other for evaluation. The training dataset is used to train the retrieval algorithm, while the evaluation dataset is used to assess the retrieval performance. The dataset of IWC profiles is derived from CloudSat reflectivity and CALIOP lidar observations. The retrieval of the ice water content IWC profiles from the computed observations is achieved in two steps. In the first step, a class, out of 18 potential classes characterized by different vertical distribution of IWC, is estimated from the observations. The 18 classes are predetermined based on the k-Means clustering algorithm. In the second step, the IWC profile is estimated using an Ensemble Kalman Smoother (EKS) algorithm that uses the estimated class as a priori information. The results of the study show that the synergy of lidar, radar, and radiometer observations is significant in the retrieval of the IWC profiles. Nevertheless, it should be mentioned that this synergy was found under idealized conditions, and additional work might be required to materialize it in practice. The inclusion of the lidar backscatter observations in the retrieval process has a larger impact on the retrieval performance than the inclusion of the radar observations. As ice clouds have a significant impact on atmospheric radiative processes, this work is relevant to ongoing efforts to reduce uncertainties in climate analyses and projections.

Mircea Grecu↗

Influence of Sub-grid-Scale Isentropic Transports on McRAS Evaluations using ARM-CART SCM Datasets

In GCM-physics evaluations with the currently available ARM-CART SCM datasets, McRAS produced very similar character of near surface errors of simulated temperature and humidity containing typically warm and moist biases near the surface and cold and dry biases aloft. We argued it must have a common cause presumably rooted in the model physics. Lack of vertical adjustment of horizontal transport was thought to be a plausible source. Clearly, debarring such a freedom would force the incoming air to diffuse into the grid-cell which would naturally bias the surface air to become warm and moist while the upper air becomes cold and dry, a characteristic feature of McRAS biases. Since, the errors were significantly larger in the two winter cases that contain potentially more intense episodes of cold and warm advective transports, it further reaffirmed our argument and provided additional motivation to introduce the corrections. When the horizontal advective transports were suitably modified to allow rising and/or sinking following isentropic pathways of subgrid scale motions, the outcome was to cool and dry (or warm and moisten) the lower (or upper) levels. Ever, crude approximations invoking such a correction reduced the temperature and humidity biases considerably. The tests were performed on all the available ARM-CART SCM cases with consistent outcome. With the isentropic corrections implemented through two different numerical approximations, virtually similar benefits were derived further confirming the robustness of our inferences. These results suggest the need for insentropic advective transport adjustment in a GCM due to subgrid scale motions.

Sud, Y. C.↗

Daily evaluation of 26 precipitation datasets using Stage-IV gauge-radar data for the CONUS

New precipitation (P) datasets are released regularly, following innovations in weather forecasting models, satellite retrieval methods, and multi-source merging techniques. Using the conterminous US as a case study, we evaluated the performance of 26 gridded (sub-)daily P datasets to obtain insight into the merit of these innovations. The evaluation was performed at a daily timescale for the period 2008–2017 using the Kling–Gupta efficiency (KGE), a performance metric combining correlation, bias, and variability. As a reference, we used the high-resolution (4 km) Stage-IV gauge-radar P dataset. Among the three KGE components, the P datasets performed worst overall in terms of correlation (related to event identification). In terms of improving KGE scores for these datasets, improved P totals (affecting the bias score) and improved distribution of P intensity (affecting the variability score) are of secondary importance. Among the 11 gauge-corrected P datasets, the best overall performance was obtained by MSWEP V2.2, underscoring the importance of applying daily gauge corrections and accounting for gauge reporting times. Several uncorrected P datasets outperformed gauge-corrected ones. Among the 15 uncorrected P datasets, the best performance was obtained by the ERA5-HRES fourth-generation reanalysis, reflecting the significant advances in earth system modeling during the last decade. The (re)analyses generally performed better in winter than in summer, while the opposite was the case for the satellite-based datasets. IMERGHH V05 performed substantially better than TMPA-3B42RT V7, attributable to the many improvements implemented in the IMERG satellite P retrieval algorithm. IMERGHH V05 outperformed ERA5-HRES in regions dominated by convective storms, while the opposite was observed in regions of complex terrain. The ERA5-EDA ensemble average exhibited higher correlations than the ERA5-HRES deterministic run, highlighting the value of ensemble modeling. The WRF regional convection-permitting climate model showed considerably more accurate P totals over the mountainous west and performed best among the uncorrected datasets in terms of variability, suggesting there is merit in using high-resolution models to obtain climatological P statistics. Our findings provide some guidance to choose the most suitable P dataset for a particular application.

Hylke E. Beck↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

Simultaneous Retrieval of Selected Optical Water Quality Indicators From Landsat-8, Sentinel-2, and Sentinel-3

Constructing multi-source satellite-derived water quality (WQ) products in inland and nearshore coastal waters from the past, present, and future missions is a long-standing challenge. Despite inherent differences in sensors’ spectral capability, spatial sampling, and radiometric performance, research efforts focused on formulating, implementing, and validating universal WQ algorithms continue to evolve. This research extends a recently developed machine-learning (ML) model, i.e., Mixture Density Networks (MDNs) (Pahlevan et al., 2020; Smith et al., 2021), to the inverse problem of simultaneously retrieving WQ indicators, including chlorophyll-a (Chla), Total Suspended Solids (TSS), and the absorption by Colored Dissolved Organic Matter at 440 nm (a cdom (440)), across a wide array of aquatic ecosystems. We use a database of in situ measurements to train and optimize MDN models developed for the relevant spectral measurements (400–800 nm) of the Operational Land Imager (OLI), MultiSpectral Instrument (MSI), and Ocean and Land Color Instrument (OLCI) aboard the Landsat-8, Sentinel-2, and Sentinel-3 missions, respectively. Our two performance assessment approaches, namely hold-out and leave-one-out, suggest significant, albeit varying degrees of improvements with respect to second-best algorithms, depending on the sensor and WQ indicator (e.g., 68%, 75%, 117% improvements based on the hold-out method for Chla, TSS, and a cdom (440), respectively from MSI-like spectra). Using these two assessment methods, we provide theoretical upper and lower bounds on model performance when evaluating similar and/or out-of-sample datasets. To evaluate multi-mission product consistency across broad spatial scales, map products are demonstrated for three near-concurrent OLI, MSI, and OLCI acquisitions. Overall, estimated TSS and a cdom (440) from these three missions are consistent within the uncertainty of the model, but Chla maps from MSI and OLCI achieve greater accuracy than those from OLI. By applying two different atmospheric correction processors to OLI and MSI images, we also conduct matchup analyses to quantify the sensitivity of the MDN model and best-practice algorithms to uncertainties in reflectance products. Our model is less or equally sensitive to these uncertainties compared to other algorithms. Recognizing their uncertainties, MDN models can be applied as a global algorithm to enable harmonized retrievals of Chla, TSS, and a cdom (440) in various aquatic ecosystems from multi-source satellite imagery. Local and/or regional ML models tuned with an apt data distribution (e.g., a subset of our dataset) should nevertheless be expected to outperform our global model.

Machine learning↗

Cold-Season Precipitation Sensitivity to Microphysical Parameterizations: Hydrologic Evaluations Leveraging Snow Lidar Datasets

Abstract Cloud microphysical processes are an important facet of atmospheric modeling, as they can control the initiation and rates of snowfall. Thus, parameterizations of these processes have important implications for modeling seasonal snow accumulation. We conduct experiments with the Weather Research and Forecasting (WRF V4.3.3) Model using three different microphysics parameterizations, including a sophisticated new scheme (ISHMAEL). Simulations are conducted for two cold seasons (2018 and 2019) centered on the Colorado Rockies’ ∼750-km 2 East River watershed. Precipitation efficiencies are quantified using a drying-ratio mass budget approach and point evaluations are performed against three NRCS SNOTEL stations. Precipitation and meteorological outputs from each are used to force a land surface model (Noah-MP) so that peak snow accumulation can be compared against airborne snow lidar products. We find that microphysical parameterization choice alone has a modest impact on total precipitation on the order of ±3% watershed-wide, and as high as 15% for certain regions, similar to other studies comparing the same parameterizations. Precipitation biases evaluated against SNOTEL are 15% ± 13%. WRF Noah-MP configurations produced snow water equivalents with good correlations with airborne lidar products at a 1-km spatial resolution: Pearson’s r values of 0.9, RMSEs between 8 and 17 cm, and percent biases of 3%–15%. Noah-MP with precipitation from the PRISM geostatistical precipitation product leads to a peak SWE underestimation of 32% in both years examined, and a weaker spatial correlation than the WRF configurations. We fall short of identifying a clearly superior microphysical parameterization but conclude that snow lidar is a valuable nontraditional indicator of model performance.

54 ENVIRONMENTAL SCIENCES↗

Dataset documenting rock core evaluation from Brady’s Hot Springs well BCH-03

Nineteen short sections of core were taken from the BCH-03 borehole at Brady’s Hot Springs, Churchill County, Nevada, between 2,945 ft and the total depth of 4,885 ft. This well is in a geothermal field. Data collected includes thin section photomicrographs, Medical CT tiff stacks, microCT tiff stacks, XRD, porosity (from thin sections, He porosimeter, microCT segmentation) and P/S wave velocities.

Brady geothermal field↗

Detection of Outliers in LiDAR Data Acquired by Multiple Platforms over Sorghum and Maize

High-resolution point cloud data acquired with a laser scanner from any platform contain random noise and outliers. Therefore, outlier detection in LiDAR data is often necessary prior to analysis. Applications in agriculture are particularly challenging, as there is typically no prior knowledge of the statistical distribution of points, plant complexity, and local point densities, which are crop-dependent. The goals of this study were first to investigate approaches to minimize the impact of outliers on LiDAR acquired over agricultural row crops, and specifically for sorghum and maize breeding experiments, by an unmanned aerial vehicle (UAV) and a wheel-based ground platform; second, to evaluate the impact of existing outliers in the datasets on leaf area index (LAI) prediction using LiDAR data. Two methods were investigated to detect and remove the outliers from the plant datasets. The first was based on surface fitting to noisy point cloud data via normal and curvature estimation in a local neighborhood. The second utilized the PointCleanNet deep learning framework. Both methods were applied to individual plants and field-based datasets. To evaluate the method, an F-score was calculated for synthetic data in the controlled conditions, and LAI, the variable being predicted, was computed both before and after outlier removal for both scenarios. Results indicate that the deep learning method for outlier detection is more robust than the geometric approach to changes in point densities, level of noise, and shapes. The prediction of LAI was also improved for the wheel-based vehicle data based on the coefficient of determination (R2) and the root mean squared error (RMSE) of the residuals before and after the removal of outliers.

36 MATERIALS SCIENCE↗

An AeroCom–AeroSat study: intercomparison of satellite AOD datasets for aerosol model evaluation

To better understand and characterize current uncertainties in the important observational constraint of climate models of aerosol optical depth (AOD), we evaluate and intercompare 14 satellite products, representing nine different retrieval algorithm families using observations from five different sensors on six different platforms. The satellite products (super-observations consisting of 1°×1° daily aggregated retrievals drawn from the years 2006, 2008 and 2010) are evaluated with AErosol RObotic NETwork (AERONET) and Maritime Aerosol Network (MAN) data. Results show that different products exhibit different regionally varying biases (both under- and overestimates) that may reach ±50 %, although a typical bias would be 15 %–25 % (depending on the product). In addition to these biases, the products exhibit random errors that can be 1.6 to 3 times as large. Most products show similar performance, although there are a few exceptions with either larger biases or larger random errors. The intercomparison of satellite products extends this analysis and provides spatial context to it. In particular, we show that aggregated satellite AOD agrees much better than the spatial coverage (often driven by cloud masks) within the 1°×1° grid cells. Up to ∼50 % of the difference between satellite AOD is attributed to cloud contamination. The diversity in AOD products shows clear spatial patterns and varies from 10 % (parts of the ocean) to 100 % (central Asia and Australia). More importantly, we show that the diversity may be used as an indication of AOD uncertainty, at least for the better performing products. This provides modellers with a global map of expected AOD uncertainty in satellite products, allows assessment of products away from AERONET sites, can provide guidance for future AERONET locations and offers suggestions for product improvements. We account for statistical and sampling noise in our analyses. Sampling noise, variations due to the evaluation of different subsets of the data, causes important changes in error metrics. The consequences of this noise term for product evaluation are discussed.

Aerosol↗

Evaluation of Global Observations-Based Evapotranspiration Datasets and IPCC AR4 Simulations

Quantification of global land evapotranspiration (ET) has long been associated with large uncertainties due to the lack of reference observations. Several recently developed products now provide the capacity to estimate ET at global scales. These products, partly based on observational data, include satellite ]based products, land surface model (LSM) simulations, atmospheric reanalysis output, estimates based on empirical upscaling of eddycovariance flux measurements, and atmospheric water balance datasets. The LandFlux-EVAL project aims to evaluate and compare these newly developed datasets. Additionally, an evaluation of IPCC AR4 global climate model (GCM) simulations is presented, providing an assessment of their capacity to reproduce flux behavior relative to the observations ]based products. Though differently constrained with observations, the analyzed reference datasets display similar large-scale ET patterns. ET from the IPCC AR4 simulations was significantly smaller than that from the other products for India (up to 1 mm/d) and parts of eastern South America, and larger in the western USA, Australia and China. The inter-product variance is lower across the IPCC AR4 simulations than across the reference datasets in several regions, which indicates that uncertainties may be underestimated in the IPCC AR4 models due to shared biases of these simulations.

Mueller, B.↗

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson↗

Simultaneous cross-evaluation of heterogeneous E. coli datasets via mechanistic simulation

The extensive heterogeneity of biological data poses challenges to analysis and interpretation. Construction of a large-scale mechanistic model of Escherichia coli enabled us to integrate and cross-evaluate a massive, heterogeneous dataset based on measurements reported by various groups over decades. We identified inconsistencies with functional consequences across the data, including that the total output of the ribosomes and RNA polymerases described by data are not sufficient for a cell to reproduce measured doubling times, that measured metabolic parameters are neither fully compatible with each other nor with overall growth, and that essential proteins are absent during the cell cycle—and the cell is robust to this absence. Finally, considering these data as a whole leads to successful predictions of new experimental outcomes, in this case protein half-lives.

59 BASIC BIOLOGICAL SCIENCES↗

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling↗

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins↗

Calculation of top-of-atmosphere, surface and atmospheric cloud radiative kernels and feedbacks based on ISCCP-H datasets

This study aims to create observation-based cloud radiative kernel (CRK) datasets and evaluate them by direct comparison of CRK and the CRK-derived cloud feedback datasets. Based on the International Satellite Cloud Climatology Project (ISCCP) H datasets, we calculate CRKs (called FH CRKs) as 2D joint function/histogram of cloud optical depth and cloud top pressure for shortwave, longwave, and their sum, Net, at the top of atmosphere (TOA), as well as, for the first time, at the surface (SFC) and in the atmosphere (ATM). The direct comparison shows that FH agrees reasonably well with three other TOA CRK datasets. With cloud fraction change (CFC) datasets of the same histogram for doubled-CO2 simulation from 10 CFMIP1 models, we derive all the TOA, SFC and ATM cloud feedback using the FH CRKs. Our TOA cloud feedback is highly similar to the previous counterparts. Based on the comparison for the 4 CRK datasets and the 10 CFC datasets, we estimate the uncertainty budget for the CRK-derived cloud feedback and show that the CFC-associated uncertainty contributes > 98.5% of the total cloud feedback uncertainty while CRK’s is very small. Our preliminary evaluation shows that some near-zero/small cloud feedback in the TOA-alone feedback indeed results from the compensation of sizable cloud feedback of the SFC and ATM feedback, demonstrating that they help reveal some significant surface and atmospheric cloud feedback whose sum appears insignificant in TOA-alone feedback

cloud radiative kernel (CRK) datasets↗

Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing

Due to the growing use of large language models (LLMs) by developers and researchers, it has become essential to reliably evaluate their ability to generate code that uses specialized libraries. We explore the use of compile-time and run-time evaluation of LLM-generated Kokkos code through extending the methods used by OpenAI with the HumanEval dataset. Our evaluation framework is based on the first 40 prompts from the Kokkos138 dataset. We start by discussing two different forms of LLM prompting, using entirely plain English or providing pseudocode for added context. These two methods are used to generate Kokkos code with the Llama-3.1-8B-Instruct and CodeQwen1.5-7B-Chat models. We found that both forms of prompting led to high failure rates and difficulties with reliably parsing LLM-generated code, while prompts with pseudocode for context generally led to improved results on more complicated tests.

97 MATHEMATICS AND COMPUTING↗