Avocado Atlas Sweep 1781365822100
Executive Summary

Executive Summary

Headline insights on price trends, regional variation, and type comparison

Observations
10000
Price_Range_Min
0.44
Price_Range_Max
3.25
Average_Price_Overall
1.4046
Total_Volume_Millions
8320.7895
Organic_Premium_Pct
42.0453
Analysis of 10,000 weekly avocado price observations across 54 regions and 2 product types spanning January 04, 2015 to March 25, 2018 reveals that average prices remained relatively stable by approximately -0.3% over the period. Prices ranged from $0.44 to $3.25, with regional variation of $0.77 between Houston (lowest) and Hartford Springfield (highest). Organic avocados commanded a premium of 42.0% compared to conventional varieties, averaging $1.65 versus $1.16 respectively.
Suggested Interpretation

Analysis of 10,000 weekly avocado price observations across 54 regions and 2 product types spanning January 04, 2015 to March 25, 2018 reveals that average prices remained relatively stable by approximately -0.3% over the period. Prices ranged from $0.44 to $3.25, with regional variation of $0.77 between Houston (lowest) and Hartford Springfield (highest). Organic avocados commanded a premium of 42.0% compared to conventional varieties, averaging $1.65 versus $1.16 respectively.

Overview

Analysis Overview

Dataset scope: observations, regions, time period, and key price metrics

Total Observations10000
Regions Analysed54
Product Types2
Price Range$0.44 – $3.25
Organic Premium42.0%
Suggested Interpretation

This analysis examined 10,000 weekly avocado price observations across 54 regions and 2 product types (conventional and organic) spanning January 04, 2015 to March 25, 2018. Average prices ranged from $0.44 to $3.25 with an overall mean of $1.40. Organic avocados commanded a premium of approximately 42.0% over conventional varieties, with total volume reaching 8320.8 million units across the period.

Data Preparation

Data Quality

Row accounting and data completeness

Initial Rows10000
Final Rows10000
Rows Removed0
Retention Rate100%
Suggested Interpretation

All 10,000 observations were retained for analysis with no rows removed. No missing values required imputation, and all price and volume data passed validation checks. Temporal coverage spans from March 25, 2018 back to January 04, 2015 across weekly reporting periods.

Visualization

Price Trend Over Time

Monthly average avocado prices over time

Suggested Interpretation

Monthly average avocado prices span 39 periods (Jan 2015 to Mar 2018). Peak price of $1.82 occurred in Sep 2017. The lowest price was $1.19 in May 2016. The first period recorded $1.35 and the final period $1.35. The market exhibits an overall relatively stable pattern with moderate fluctuations.

Visualization

Conventional vs. Organic Price Distribution

Box plot comparing price distributions and central tendencies between conventional and organic avocados

Suggested Interpretation

Organic avocados carry a premium of approximately 42.0% relative to conventional varieties. Conventional avocados average $1.16 with a median of $1.13 and an interquartile range of $0.34. Organic avocados average $1.65 with a median of $1.62 and an interquartile range of $0.44. The comparable spread in both distributions—ranging from $1.76 (conventional) to $2.81 (organic)—indicates stable pricing consistency within each category.

Visualization

Average Price by Region

Regional variation in average avocado prices. Shows all regions ranked by average price, with top 12 distinct regions plus 'Other' category.

Suggested Interpretation

Regional avocado prices vary substantially across the 54 regions analyzed. Hartford Springfield commands the highest average price at $1.82, while Houston has the lowest at $1.05. The spread between highest and lowest is $0.77, reflecting significant regional pricing variation. The average price across all regions is $1.41.

Data Table

Regional Price and Volume Summary

Summary of regional market size, pricing, and sampling depth

RegionTotal VolumeAverage PriceObservation Count
Total US3.062e+091.306181
West5.698e+081.298195
California5.556e+081.423191
South Central5.52e+081.105195
Northeast3.682e+081.601179
Southeast3.52e+081.387184
Great Lakes3.289e+081.35193
Los Angeles2.975e+081.195181
Midsouth2.802e+081.404191
Plains1.637e+081.458185
Dallas Ft Worth1.358e+081.085211
New York1.309e+081.702177
Houston1.049e+081.054186
Phoenix Tucson1.012e+081.242179
West Tex New Mexico7.965e+071.261188
San Francisco7.46e+071.822189
Denver7.267e+071.235188
Baltimore Washington7.265e+071.53189
Chicago7.031e+071.571186
Portland6.743e+071.312196
Boston5.735e+071.531201
Seattle5.064e+071.442161
Atlanta4.968e+071.33184
San Diego4.592e+071.433183
Northern New England4.468e+071.446193
Miami Ft Lauderdale4.466e+071.44170
Sacramento4.379e+071.616184
Philadelphia4.256e+071.576172
Orlando3.637e+071.482191
Detroit3.521e+071.27190
South Carolina3.296e+071.406191
Tampa3.272e+071.414176
Raleigh Greensboro2.993e+071.529201
Las Vegas2.674e+071.408179
Hartford Springfield2.46e+071.824163
New Orleans Mobile2.362e+071.308173
Harrisburg Scranton2.353e+071.513188
Richmond Norfolk2.259e+071.288174
Cincinnati Dayton2.11e+071.195171
Nashville1.962e+071.186179
Charlotte1.825e+071.644191
St Louis1.795e+071.473195
Grand Rapids170842051.474187
Indianapolis1.586e+071.306186
Columbus1.562e+071.28184
Jacksonville1.522e+071.519186
Roanoke1.226e+071.278185
Buffalo Rochester1.173e+071.495180
Pittsburgh9.68e+061.363178
Louisville8.463e+061.311202
Albany8.417e+061.55181
Spokane8.204e+061.428183
Boise7.889e+061.308178
Syracuse6.223e+061.519196
Suggested Interpretation

The 54 regions in this analysis show substantial variation in market size and pricing. Total US dominates by volume with 3,061,709,672 units—approximately 36.8% of the total market across all regions. Regional average prices span from $1.05 in Houston to $1.82 in Hartford Springfield, a spread of $0.77 that reflects meaningful pricing tier variation across geographies. Observation depth per region averages 185 weekly samples, with 10,000 total observations distributed across all regions to support consistent regional estimates.

Methodology

Methodology

Statistical methodology and diagnostics for Avocado Atlas Sweep

Statistical Method

Avocado Atlas Sweep

Examines avocado prices across a multi-year period spanning multiple regions and type categories (conventional vs. organic). Traces temporal price evolution using weekly data, identifies regional pricing differences via regional aggregation, and compares price distributions by type to surface pricing dynamics.

Software & Citation
MCP Analytics · mcpanalytics.ai
Code Appendix

Analysis Code

Complete R source code for this analysis

Avocado Atlas Sweep

Examines avocado prices across a multi-year period spanning multiple regions and type categories (conventional vs. organic). Traces temporal price evolution using weekly data, identifies regional pricing differences via regional aggregation, and compares price distributions by type to surface pricing dynamics.

Why This Method?

Descriptive aggregation and temporal analysis provide a comprehensive view of avocado market structure without requiring statistical modeling. This approach reveals pricing trends, regional variation, and type-based differences that inform market positioning and sourcing decisions.

What This Analysis Covers

  • Price Trends Over Time: Tracks average price evolution across the analysis period to detect patterns
  • Conventional vs Organic: Compares price distributions and identifies quality/type premiums
  • Regional Price Variation: Identifies which regions command highest/lowest prices and market concentration
  • Regional Market Summary: Summarizes volume, pricing, and observation counts by geography
suppressPackageStartupMessages(library(htmltools))
suppressPackageStartupMessages(library(jsonlite))
suppressPackageStartupMessages(library(plotly))
suppressPackageStartupMessages(library(DT))
suppressPackageStartupMessages(library(htmlwidgets))
suppressPackageStartupMessages(library(arrow))
suppressPackageStartupMessages(library(knitr))
suppressPackageStartupMessages(library(rmarkdown))
suppressPackageStartupMessages(library(dplyr))
suppressPackageStartupMessages(library(tidyr))
suppressPackageStartupMessages(library(ggplot2))
suppressPackageStartupMessages(library(stringr))
suppressPackageStartupMessages(library(lubridate))
suppressPackageStartupMessages(library(broom))
suppressPackageStartupMessages(library(survival))
suppressPackageStartupMessages(library(Matrix))
suppressPackageStartupMessages(library(cluster))
suppressPackageStartupMessages(library(data.table))

Step 1: Data Preparation

Row accounting (LWS-74) — no filtering, all rows analyzed.

initial_rows <- nrow(df)
  final_rows   <- nrow(df)
  rows_removed <- 0L

Step 2: Temporal Aggregation

Create monthly aggregations with proper chronological ordering for line chart. Use lubridate::ymd() to parse dates, then monthly buckets with YYYYMM sort key.

df$date_parsed <- lubridate::ymd(df$date)
  df$year_month  <- lubridate::floor_date(df$date_parsed, "month")
  df$yyyymm      <- format(df$year_month, "%Y%m")
  df$month_label <- format(df$year_month, "%b %Y")

  price_over_time <- df %>%
    dplyr::group_by(yyyymm, month_label) %>%
    dplyr::summarise(
      average_price = mean(average_price, na.rm = TRUE),
      n = dplyr::n(),
      .groups = "drop"
    ) %>%
    dplyr::filter(n >= 5) %>%
    dplyr::arrange(yyyymm) %>%
    dplyr::select(date_period = month_label, average_price)

Step 3: Price by Type (Box Plot)

Humanize type names (conventional/organic → capitalized).

price_by_type <- df %>%
    dplyr::mutate(type = tools::toTitleCase(type)) %>%
    dplyr::select(type, average_price)

Step 4: Price by Region (Horizontal Bar, Top-12 Rollup)

Split CamelCase region names, aggregate by region, sort by price descending, cap at 12 regions with "Other" rollup.

df$region_display <- gsub("([a-z])([A-Z])", "\\1 \\2", df$region)

  price_by_region_all <- df %>%
    dplyr::group_by(region_display) %>%
    dplyr::summarise(average_price = mean(average_price, na.rm = TRUE), .groups = "drop") %>%
    dplyr::arrange(dplyr::desc(average_price))

  n_regions <- nrow(price_by_region_all)
  if (n_regions > 12) {
    price_by_region <- dplyr::bind_rows(
      dplyr::slice(price_by_region_all, 1:12),
      data.frame(
        region_display = paste0("Other(", n_regions - 12, ")"),
        average_price = mean(price_by_region_all$average_price[13:n_regions])
      )
    )
  } else {
    price_by_region <- price_by_region_all
  }

Step 5: Regional Summary (Table)

Aggregate volume, price, and observation count by region.

region_summary <- df %>%
    dplyr::group_by(region_display) %>%
    dplyr::summarise(
      total_volume = sum(total_volume, na.rm = TRUE),
      average_price = mean(average_price, na.rm = TRUE),
      observation_count = dplyr::n(),
      .groups = "drop"
    ) %>%
    dplyr::arrange(dplyr::desc(total_volume))

Step 7: Template Variables for Prose

Format size strings for Rule T (template-based narratives).

tpl_vars <- list(
    n_observations_fmt = format(final_rows, big.mark = ","),
    n_regions = n_regions,
    n_types = length(unique(df$type)),
    date_range = paste0(
      format(min(df$date_parsed, na.rm = TRUE), "%B %d, %Y"),
      " to ",
      format(max(df$date_parsed, na.rm = TRUE), "%B %d, %Y")
    ),
    price_min_str = sprintf("$%.2f", price_range_min),
    price_max_str = sprintf("$%.2f", price_range_max),
    price_avg_str = sprintf("$%.2f", mean(df$average_price, na.rm = TRUE)),
    organic_premium_str = sprintf("%.1f%%", if (!is.na(organic_premium_pct)) organic_premium_pct else 0),
    organic_avg_str = sprintf("$%.2f", if (length(organic_price) > 0) organic_price else 0),
    conventional_avg_str = sprintf("$%.2f", if (length(conventional_price) > 0) conventional_price else 0)
  )

Compute shared resources

shared <- compute_shared(df, params)

Finalize (do not modify)

Narrative: Peak, Trough, Trend, Currency Formatting

n_months <- nrow(pot)

  # Format currency helper (base-R, no scales package)
  fmt_currency <- function(x) {
    paste0("$", formatC(x, format = "f", digits = 2, big.mark = ","))
  }

  if (n_months > 0) {
    first_price <- pot$average_price[1]
    last_price <- pot$average_price[n_months]
    peak_idx <- which.max(pot$average_price)
    peak_price <- pot$average_price[peak_idx]
    peak_month <- pot$date_period_label[peak_idx]
    trough_idx <- which.min(pot$average_price)
    trough_price <- pot$average_price[trough_idx]
    trough_month <- pot$date_period_label[trough_idx]

    displayed_range <- paste0(
      pot$date_period_label[1], " to ",
      pot$date_period_label[n_months]
    )

    # Trend direction: compare first 3 vs last 3 months
    if (n_months >= 3) {
      first3_mean <- mean(pot$average_price[1:min(3, n_months)], na.rm = TRUE)
      last3_mean <- mean(pot$average_price[max(1, n_months - 2):n_months], na.rm = TRUE)
      if (last3_mean > first3_mean * 1.1) {
        trend <- "upward trend, reflecting rising avocado prices."
      } else if (last3_mean < first3_mean * 0.9) {
        trend <- "downward trend, indicating price decreases."
      } else {
        trend <- "relatively stable pattern with moderate fluctuations."
      }
    } else {
      trend <- "limited data for trend assessment."
    }
  } else {
    first_price <- last_price <- peak_price <- trough_price <- 0
    peak_month <- trough_month <- displayed_range <- "N/A"
    trend <- "no data available."
  }

  # Build narrative text
  text <- paste0(
    "Monthly average avocado prices span ", n_months, " periods(",
    displayed_range, "). Peak price of ", fmt_currency(peak_price),
    " occurred in ", peak_month, ". The lowest price was ",
    fmt_currency(trough_price), " in ", trough_month, ". ",
    "The first period recorded ", fmt_currency(first_price),
    " and the final period ", fmt_currency(last_price), ". ",
    "The market exhibits an overall ", trend
  )

  # Append capping note if applied
  if (capped) {
    text <- paste0(
      text, " (Displayed prices capped at q99 = ",
      fmt_currency(cap_value), " for clarity.)"
    )
  }

  # Prepare output data frame with ISO-8601 date_period and capped average_price
  data_out <- data.frame(
    date_period = pot$date_period,
    average_price = pot$average_price,
    stringsAsFactors = FALSE
  )

  list(
    title = "Price Trend Over Time",
    description = "Monthly average avocado prices over time",
    text = text,
    columns = list(
      date_period = list(role = "temporal"),
      average_price = list(role = "value", format = "currency", symbol = "$")
    ),
    data = list(price_over_time = data_out)
  )
}

# Card: regional_price_mix (horizontal_bar)
# Returns: list(title, description, text, data)

card_regional_price_mix <- function(shared, df, params) {
  # Extract and rename column to match spec
  card_df <- shared$price_by_region %>%
    dplyr::rename(region = region_display)

  # Compute narrative stats from ALL regions (not just displayed top-12)
  # to avoid claiming "Other" aggregation is the lowest individual region
  all_region_stats <- df %>%
    dplyr::mutate(region = gsub("([a-z])([A-Z])", "\\1 \\2", region)) %>%
    dplyr::group_by(region) %>%
    dplyr::summarise(average_price = mean(average_price, na.rm = TRUE), .groups = "drop") %>%
    dplyr::arrange(dplyr::desc(average_price))

  n_regions <- nrow(all_region_stats)

  if (n_regions > 0) {
    # Highest from all regions
    highest_region <- all_region_stats$region[1]
    highest_price <- all_region_stats$average_price[1]
    # Lowest from all regions (true minimum, not "Other" aggregation)
    lowest_region <- all_region_stats$region[n_regions]
    lowest_price <- all_region_stats$average_price[n_regions]
    price_spread <- highest_price - lowest_price
    avg_price <- mean(all_region_stats$average_price, na.rm = TRUE)
  } else {
    highest_region <- "N/A"
    highest_price <- 0
    lowest_region <- "N/A"
    lowest_price <- 0
    price_spread <- 0
    avg_price <- 0
  }

  # Check for heavy-tailed outliers (monetary_card rule 1)
  display_note <- ""
  if (n_regions > 0) {
    p99_price <- quantile(card_df$average_price, 0.99, na.rm = TRUE)
    max_price <- max(card_df$average_price, na.rm = TRUE)

    if (max_price > p99_price && (max_price / p99_price) >= 3) {
      # Cap for display at p99
      card_df$average_price <- pmin(card_df$average_price, p99_price)
      display_note <- paste0(" (capped at 99th percentile: $", round(p99_price, 2), " for display)")
    }
  }

  # LAT-59 #1 (categorical_axis #6): For horizontal_bar ranking card,
  # sort ASCENDING so highest bars render at TOP
  card_df <- card_df %>%
    dplyr::arrange(average_price)

  # Format currency (monetary_card rule 3)
  fmt_dollar <- function(x) {
    paste0("$", formatC(x, format = "f", big.mark = ",", digits = 2))
  }

  # Build narrative from actual computed values
  narrative_text <- paste0(
    "Regional avocado prices vary substantially across the ",
    n_regions,
    " regions analyzed. ",
    highest_region,
    " commands the highest average price at ",
    fmt_dollar(highest_price),
    ", while ",
    lowest_region,
    " has the lowest at ",
    fmt_dollar(lowest_price),
    ". The spread between highest and lowest is ",
    fmt_dollar(price_spread),
    ", reflecting significant regional pricing variation. ",
    "The average price across all regions is ",
    fmt_dollar(avg_price),
    ".",
    display_note
  )

  list(
    title = "Average Price by Region",
    description = paste0(
      "Regional variation in average avocado prices. ",
      "Shows all regions ranked by average price, with top 12 distinct regions plus &#x27;Other' category."
    ),
    text = narrative_text,
    data = list(
      price_by_region = card_df
    )
  )
}

# Card: regional_summary (table)
# Returns: list(title, description, text, data)
# Displays regional aggregation: volume, average price, observation count per region

card_regional_summary <- function(shared, df, params) {
  # Extract pre-computed regional summary table (sorted by volume descending)
  region_data <- shared$region_summary

  # Validate data availability
  if (is.null(region_data) || nrow(region_data) == 0) {
    return(list(
      title = "Regional Price and Volume Summary",
      description = "Summary of regional market size, pricing, and sampling depth",
      text = "Insufficient data to compute regional summary.",
      data = list()
    ))
  }

  # Compute statistics for narrative
  n_regions <- nrow(region_data)
  top_region <- region_data$region_display[1]
  top_volume <- region_data$total_volume[1]

  # Price and observation statistics across all regions
  price_max_region <- region_data$region_display[which.max(region_data$average_price)]
  price_max <- max(region_data$average_price, na.rm = TRUE)
  price_min_region <- region_data$region_display[which.min(region_data$average_price)]
  price_min <- min(region_data$average_price, na.rm = TRUE)
  price_spread <- price_max - price_min

  total_obs <- sum(region_data$observation_count, na.rm = TRUE)
  avg_obs_per_region <- round(mean(region_data$observation_count, na.rm = TRUE), 0)

  # Volume concentration: what % of total is top region?
  total_volume <- sum(region_data$total_volume, na.rm = TRUE)
  top_volume_pct <- round(100 * top_volume / total_volume, 1)

  # Build narrative addressing insight focus: market size, pricing tiers, observation depth
  regional_summary_text <- paste0(
    "The ",
    n_regions,
    " regions in this analysis show substantial variation in market size and pricing. ",
    top_region,
    " dominates by volume with ",
    format(round(top_volume, 0), big.mark = ","),
    " units—approximately ",
    top_volume_pct,
    "% of the total market across all regions. ",
    "Regional average prices span from ",
    sprintf("$%.2f", price_min),
    " in ",
    price_min_region,
    " to ",
    sprintf("$%.2f", price_max),
    " in ",
    price_max_region,
    ", a spread of ",
    sprintf("$%.2f", price_spread),
    " that reflects meaningful pricing tier variation across geographies. ",
    "Observation depth per region averages ",
    format(avg_obs_per_region, big.mark = ","),
    " weekly samples, with ",
    format(total_obs, big.mark = ","),
    " total observations distributed across all regions to support consistent regional estimates."
  )

  # Prepare display table: rename columns for readability
  display_table <- region_data %>%
    dplyr::rename(
      "Region" = region_display,
      "Total Volume" = total_volume,
      "Average Price" = average_price,
      "Observation Count" = observation_count
    )

  list(
    title = "Regional Price and Volume Summary",
    description = "Summary of regional market size, pricing, and sampling depth",
    text = regional_summary_text,
    data = list(region_summary = display_table)
  )
}

# Card: tldr (tldr)
# Returns: tldr -> list(title, description, metrics, text);
#          other -> list(title, description, text, data)

card_tldr <- function(shared, df, params) {
  tv <- shared$tpl_vars
  metrics <- shared$metrics
  price_over_time <- shared$price_over_time
  price_by_type <- shared$price_by_type

  # Compute price trend direction
  if (nrow(price_over_time) >= 2) {
    first_price <- price_over_time$average_price[1]
    last_price <- price_over_time$average_price[nrow(price_over_time)]
    price_change <- last_price - first_price
    trend_direction <- if (price_change > 0.1) "increased" else if (price_change < -0.1) "decreased" else "remained relatively stable"
    price_change_pct <- round(100 * price_change / first_price, 1)
  } else {
    trend_direction <- "stable"
    price_change_pct <- 0
  }

  # Compute regional price variance (max - min) from all individual regions (not rolled-up "Other")
  # Match the regional_summary card's use of all actual regions, not the top-12+Other aggregation
  all_region_stats <- df %>%
    dplyr::mutate(region = gsub("([a-z])([A-Z])", "\\1 \\2", region)) %>%
    dplyr::group_by(region) %>%
    dplyr::summarise(average_price = mean(average_price, na.rm = TRUE), .groups = "drop") %>%
    dplyr::arrange(dplyr::desc(average_price))

  if (nrow(all_region_stats) >= 2) {
    regional_max <- max(all_region_stats$average_price, na.rm = TRUE)
    regional_min <- min(all_region_stats$average_price, na.rm = TRUE)
    regional_spread <- regional_max - regional_min
    highest_region <- all_region_stats$region[1]
    lowest_region <- all_region_stats$region[nrow(all_region_stats)]
  } else {
    regional_spread <- 0
    highest_region <- "N/A"
    lowest_region <- "N/A"
  }

  # Compute type comparison
  type_stats <- price_by_type %>%
    dplyr::group_by(type) %>%
    dplyr::summarise(
      mean_price = mean(average_price, na.rm = TRUE),
      count = dplyr::n(),
      .groups = "drop"
    ) %>%
    dplyr::arrange(dplyr::desc(mean_price))

  organic_premium_pct <- metrics$Organic_Premium_Pct %||% 0

  # Assemble narrative using computed values only
  tldr_text <- paste0(
    "Analysis of ",
    tv$n_observations_fmt,
    " weekly avocado price observations across ",
    tv$n_regions,
    " regions and ",
    tv$n_types,
    " product types spanning ",
    tv$date_range,
    " reveals that average prices ",
    trend_direction,
    " by approximately ",
    sprintf("%.1f%%", price_change_pct),
    " over the period. Prices ranged from ",
    tv$price_min_str,
    " to ",
    tv$price_max_str,
    ", with regional variation of ",
    sprintf("$%.2f", regional_spread),
    " between ",
    lowest_region,
    " (lowest) and ",
    highest_region,
    " (highest). Organic avocados commanded a premium of ",
    tv$organic_premium_str,
    " compared to conventional varieties, averaging ",
    tv$organic_avg_str,
    " versus ",
    tv$conventional_avg_str,
    " respectively."
  )

  list(
    title = "Executive Summary",
    description = "Headline insights on price trends, regional variation, and type comparison",
    metrics = metrics,
    text = tldr_text
  )
}
Your data has more stories to tell. Run any analysis on your own data — validated R modules, interactive reports, AI insights, and PDF export. 500 free credits on signup.
Try Free — No Signup Sign Up Free

Report an Issue

Tell us what's wrong. You'll get a free re-run of this analysis so you can try again with different parameters. If the re-run still doesn't meet your expectations, we'll refund your credits.

Want to run this analysis on your own data? Upload CSV — Free Analysis See Pricing