AdithyaSK/data_agent_rl_environment_train_subset_100
data agent rl environment train subset 100: Harbor dataset on Hugging Face with 100 tasks. A 100-task quick-iteration subset of the data-agent RL training suite. All tasks are L1 difficulty (the easiest tier) with a numeric reward function — chosen so RL/eval loops converge fast and grade…
Tasks
- What is the total number of police death cases recorded in the dataset?
- How many hospitals in the dataset have proprietary ownership?
- What is the highest price listed for any product in the dataset?
- How many columns are present in the dataset after removing the 'Timestamp' column?
- What is the average number of comments per entry in the dataset?
- What is the difference between the highest and lowest opening prices in the dataset?
- What is the maximum closing price recorded in the Bitcoin price dataset?
- What is the difference in ROC AUC scores between the optimized SVM model (0.577) and the optimized Random Forest model (0.692) on the test…
- How many missing values were present in the 'type2' column before imputation?
- What is the average number of total victims per mass shooting incident in the dataset?
- How many unique cereal manufacturers are represented in the dataset?
- What is the critical value of the chi-square distribution at 95% confidence level with 12 degrees of freedom used to determine statistical…
- How many numeric columns are present in the dataset before additional processing?
- What is the range of the health rating metric (difference between maximum and minimum values) in the dataset?
- What is the median rating value for cereals in the dataset?
- How many missing values were present in the 'Item Weight' column of the training dataset before imputation?
- What is the interquartile range (IQR) for the SepalWidthCm feature?
- What is the correlation coefficient between Sepal Length and Petal Length in the dataset?
- What is the total number of bombing missions recorded in the dataset?
- What is the highest OvertimePay amount recorded in the dataset?
- What is the correlation coefficient between Glucose levels and Outcome in the dataset?
- How many distinct years are covered in the global religious population dataset?
- How many features in the dataset contain missing values according to the null value analysis?
- What is the maximum observed value of the 'K' chemical composition feature in the glass samples?
- How many columns in the training dataset contain missing values that were identified during preprocessing?
- What is the highest single-season passing yardage achieved by a player in the dataset?
- What is the total number of drug-related deaths recorded in the dataset?
- What is the average number of bedrooms in the houses included in the dataset?
- How many distinct fruit categories are present in the dataset according to the unique values in the fruit label column?
- What is the maximum quantity ordered in the dataset?
- What is the maximum price of a diamond in the dataset before removing rows with zero dimensions?
- What is the total number of datasets in the original Kaggle dataset collection analyzed in this study?
- What is the mean pixel intensity value of the first pixel (pixel1) across all images in the training dataset?
- What is the median value of the 'concave points worst' feature across all samples in the dataset?
- What is the average base total statistic across all Pokémon in the dataset?
- What is the 75th percentile value for the special attack (sp attack) attribute among all Pokémon?
- How many non-null entries exist in the "Unnamed: 2" column of the original dataset?
- How many missing values are present in the 'event type' column of the dataset?
- What is the minimum value of 'fixed acidity' in the dataset?
- How many devices in the dataset have a price range classification of 3?
- What is the most common number of cylinders in the original dataset before outlier removal?
- How many unique categories are present in the PaymentMethod column of the dataset?
- What is the total number of hourly records in the Beijing dataset before removing missing values?
- What is the average tenure (in months) of customers in the dataset?
- What is the median age of all patients in the dataset?
- How many features are included in the model after removing 'animal name' and 'class type'?
- How many features remain in the dataset after removing 'id', 'diagnosis', and 'Unnamed: 32' columns?
- What is the correlation between Speed and Special Attack (Sp. Atk) in the dataset?
- How many numerical features are used as input variables for the machine learning models in this analysis?
- How many different stock symbols are present in the dataset?
- How many missing values were present in the original training dataset before the NaN value cleaning process?
- Which wine quality rating is the most frequent in the dataset?
- What is the average body mass index (BMI) of patients in the dataset?
- What was the exact number of missing values in the MINIMUM PAYMENTS column before imputation?
- What is the total number of unique values present in the 'stalk-root' feature, including the missing value indicator?
- What is the population of the most populous district in the dataset?
- Which floor value (e.g., 1.0, 2.0, etc.) has the highest frequency in the dataset?
- How many unique classes are present in the dataset based on the label distribution?
- What is the mean Heating Load in the building energy efficiency dataset?
- What is the maximum UnitPrice recorded in the dataset?
- What is the highest frequency of any wine quality rating in the original dataset?
- What is the mean of the total length (Petal Length + Sepal Length) across all samples?
- What is the correlation coefficient between Glucose levels and diabetes diagnosis (Outcome) in the original dataset?
- What is the average number of open accounts held by customers in the dataset?
- What is the most common outcome (0=No diabetes, 1=Diabetes) in the original dataset?
- What is the kurtosis value for the Petal Length feature in the dataset?
- How many distinct quality categories exist in the wine quality dataset?
- What is the highest Speed value recorded in the Pokémon dataset?
- What is the median living area (sqft living) of the houses in the dataset?
- What is the standard deviation of the age distribution in the dataset?
- What is the correlation coefficient between carat weight and diamond price in the original dataset before any transformations?
- What is the skewness value of the y feature (diamond height in mm) in the original dataset?
- What is the median North American sales value in the dataset?
- What is the covariance between Sepal Length and Petal Width in the raw (non-standardized) iris dataset?
- What is the median global sales value across all games in the dataset?
- What is the median North American sales value across all games in the dataset?
- What is the highest recorded final grade (G3) achieved by any student in the dataset?
- What is the average (mean) year of release for video games in this dataset?
- How many missing values were present in the 'open' column before data cleaning?
- How many unique categorical values existed in the Item Type column before label encoding?
- What was the mean value used to fill missing entries in the 'total bedrooms' column before applying random choice sampling for imputation?
- How many video games in the dataset originally contained missing values in the 'Year' column before handling them with fillna()?
- How many patients were removed from the dataset due to invalid age entries (age -1 and 115)?
- What is the median North American sales value for all video games in the dataset?
- What is the median speed value across all Pokémon in the dataset?
- How many entries in the dataset have missing values in the 'Year' column?
- How many video games are categorized as "SuperHit" based on global sales (≥60 million)?
- What is the average residual sugar content of the wines in the dataset?
- What is the total number of video game entries recorded in the dataset?
- What was the number of missing values present in the 'total bedrooms' column before data cleaning?
- How many entries in the dataset contain missing values in the "Year" column?
- How many unique categories are present in the 'safety' feature before encoding?
- What is the total number of samples in the Boston housing dataset?
- What is the median value of Sepal Length (cm) across all samples in the dataset?
- How many passengers had missing values in the Age column before imputation?
- What is the maximum global sales value recorded for any single video game in the dataset?
- How many customers in the dataset made a purchase?
- How many entries in the dataset have missing values in the 'Year' column?
- What is the correlation coefficient between sales and quantity ordered in the dataset?
- What is the total number of missing values across all columns in the dataset?
AdithyaSK/data_agent_rl_environment_train_subset_100 on the Hugging Face Hub