Analytics case study

Online News Performance Analysis

A decision-focused investigation of what drives article sharing—and what the data cannot reliably promise.

Role
Sole analyst
Tools
Power Query, Power BI
Dataset
40,000 articles
Focus
Editorial performance

The brief

The project examined the UCI Online News Popularity dataset, originally collected from Mashable, to understand patterns associated with article sharing. The objective was not to manufacture a formula for virality. It was to give an editorial team a clearer view of performance, timing, content categories, and the limits of prediction.

40Karticles analysed
135Mtotal shares
1.4Kmedian shares
61variables reviewed
Power BI overview with article count, total shares, median shares and performance distributions
Executive overview: headline metrics and the shape of article-sharing performance.Open full-size image ↗

Approach

I used Power Query to inspect, clean, and prepare the data, then built a six-page Power BI report around the questions an editorial team would actually ask:

  • Which content categories generate the strongest typical performance?
  • How does publishing day relate to sharing?
  • Which article attributes appear related to reach?
  • Are any individual features strong enough to predict success on their own?

Because share counts are strongly skewed by exceptional articles, the analysis pairs total shares with medians and distributions. That keeps high-volume outliers from becoming the whole story.

Dashboard comparing online news performance by content category
Category performance compared using both volume and typical shares.Open full-size image ↗

What the evidence showed

  • Social-media articles had the highest median performance at about 2.1K shares.
  • Saturday and Sunday articles recorded medians of about 2K and 1.9K shares, respectively, despite lower publishing volume.
  • The overall median of 1.4K shares is more representative of a typical article than the mean in this long-tailed dataset.
  • No single article feature showed a strong relationship with shares. Reach appears to emerge from several interacting editorial and audience conditions.

Interpretation: timing and category are useful planning signals, not guarantees. Editorial experimentation should treat them as hypotheses to test alongside topic quality, distribution, and audience context.

Dashboard showing online news performance by publishing day and time patterns
Publishing-pattern view used to compare weekday and weekend performance.Open full-size image ↗
Dashboard examining correlations between article features and shares
Feature review: useful associations were modest, reinforcing a multi-factor interpretation.Open full-size image ↗

From insight to action

The final report recommends a test-and-learn publishing strategy: protect weekend slots for promising content, continue building social-media coverage, and test combinations of timing, channel, headline, and topic rather than optimising one variable in isolation. A lightweight editorial scorecard could track median shares, consistency, and category mix over time.

The work also makes uncertainty visible. This is an observational dataset, so the patterns support prioritisation and experiments—not causal claims.