Portfolio Project 1

Real or fake? Interrogating a dataset

Overview

If you’ve done a “find your own data” project in the past, you are probably familiar with Kaggle. But, a surprising amount of the data on Kaggle isn’t real. Some of it is openly synthetic (generated to look like a real dataset), some is simulated from a textbook or a model, some is a real dataset that has been quietly “cleaned” beyond recognition, and some was truly collected in the real world and posted online. It’s easy to get several hours into an analysis before you realize the patterns you’re describing don’t actually exist in the real world.

In this project, you’ll build up your “fake data” alarm bell. You’ll pick one dataset, investigate how it was really produced, and make a data-driven argument about whether it reflects a genuine data-collection process or was fabricated, simulated, or synthetically generated.

The verdict is not the point of the project. A carefulreport that concludes “I found strong but not conclusive signs this is synthetic” or “this is real, but it has serious data-quality problems that would have bitten me later” earns full marks if the investigation backs up those statements. What you’re being graded on is whether you can use ggplot2 and basic wrangling to interrogate a dataset and reason about ity.

What your report should contain

Keep it short: about 3 pages if rendered to PDF (including figures).

  1. Introduction one paragraph: what the dataset is, where on Kaggle you found it (link), what it claims to be, and what its data card / description says about where it came from.
  2. At least three lines of inquiry. Each one is a short subsection with:
    • the question you’re asking of the data (“do the ages cluster on round numbers?”, “should these two columns agree, and do they?”),
    • a single graph or summary table that addresses it,
    • two or three sentences on what it shows and how it moves your thinking.
  3. Verdict your best judgment, stated with a confidence level (e.g. “likely synthetic”, “probably real but poorly documented”, “genuinely can’t tell, and here’s why”), justified by the evidence above.
  4. Reflection one short paragraph: if you’d found this out on day one, how would it change what you’d do with this dataset next? Or: what would you need from the dataset’s authors to be confident?

Your rendered file should not echo the code, but the code in your .qmd should be clean and readable.

Lines of inquiry to try

You don’t need all of these, and a good project usually goes deep on a few rather than shallow on all.

  • Distributions that are too clean: suspiciously uniform, perfectly bell-shaped, no outliers, constant variance.
  • Impossible or contradictory records: a 0 where a 0 can’t happen (zero blood pressure, zero price), an end date before a start date, a total that doesn’t equal the sum of its parts, a city in the wrong state.
  • Duplicates and near-duplicates: exact repeated rows, or rows that differ only in an ID. Real collection produces some; a lot is a red flag.
  • Round-number clustering: humans and generators both love round numbers, but real instruments usually don’t.
  • Decimal precision: every value to exactly two decimals, or a column of “measurements” with more precision than any real instrument would give.
  • Correlation structure: variables that should be related in the real world (height and weight, price and size) but aren’t, or vice versa. Extremely strong correlation in variables that are surprising.
  • Categorical balance: categories split almost exactly evenly, or a grouping variable with implausibly equal group sizes.
  • Time structure: timestamps at perfectly regular intervals, no weekday/weekend effect, no gaps.
  • External benchmarks: does the mean / range / proportion match what you can look up about the real world?
  • The paper trail: the data card, the uploader’s other datasets, a cited source you can go check, a “generated with” note in the description.

Dataset shortlist

Pick one of these, or bring your own (see below). Download the CSV from Kaggle and commit it to the data/ folder of your project repo.

Your report should not just repeat what someone else concluded. If you find a discussion of the dataset’s authenticity online, cite it, but the analysis must be your own. (But I would prefer you don’t go searching for this!)

  1. Medical Cost Personal Datasets: https://www.kaggle.com/datasets/mirichoi0218/insurance.
  2. Students Performance in Exams:
    https://www.kaggle.com/datasets/spscientist/students-performance-in-exams
  3. Greenhouse Plant Growth: https://www.kaggle.com/datasets/adilshamim8/greenhouse-plant-growth-metrics
  4. Spotify Tracks Dataset:
    https://www.kaggle.com/datasets/maharshipandya/-spotify-tracks-dataset

Bring your own

If none of these speak to you, choose your own Kaggle dataset! Some general guidelines:

  • at least 6 variables, with at least 2 quantitative and 2 categorical,
  • at least 200 rows,
  • provenance that is genuinely ambiguous: do not pick a dataset whose description already tells you it’s synthetic, and don’t pick something so famous the answer is common knowledge (like penguins or titanic)
  • run your choice by me on Ed before you start so I know the dataset.

Submission

Work in a quarto file. Render to gfm, pdf, or html. I’ll distribute a GitHub skeleton repo; fill it in as you work, commit your final .qmd and rendered output, and link the repo to Gradescope. Give your report an informative title.

Rubric

A successful project will:

    • There should be at least two commits with substantial changes between them
    • Very few grammatical mistakes, spelling mistakes, or typos
    • Informative report title; readable theme; appropriate labels and font sizes
    • No package-loading messages, warnings, or other clutter in the output

An excellent project will meet all of the requirements for a successful project, plus

NoteOn not being sure

For some datasets you will not be able to prove anything. That’s realistic, and it’s fine. Grade yourself on whether your evidence is varied and correctly read and your argument is honest about its own limits. And remember only 2 of the 3 portfolio projects need to hit “excellent” for an A, so it’s OK to decide a given dataset isn’t giving you enough to work with and say so.

Can I work with someone?

This portfolio project can be done individually or with a partner. If you’d like to work with a partner, let me know by Friday so I can update the permissions on a shared repo. You may brainstorm, get feedback on graphs, and get conceptual debugging help from others; but the analysis and writing must be your own. Please cite any resources you use (title, author, link, and one line on what you used it for).

A note/reminder on AI: Large-language models (e.g. ChatGPT, Gemini, etc.) should only be used for coding or debugging help after you’ve attempted to solve the problem on your own. You should never copy and paste any course materials into a large-language model, and you should never copy and paste anything out of a large-language model into your course materials. Copying, paraphrasing, summarizing, or submitting work generated by anyone but yourself without proper attribution is considered academic dishonesty (this includes output from LLMs). You are not allowed to upload datasets or assignment prompts into a large language model.

FAQ

If you have any questions, please post them in the “Portfolio Project 1” thread on Ed.