14 Comments
User's avatar
Mata Haggis-Burridge's avatar

Fascinating work. Thank you for sharing this. I have been working on diversity in video games and agree with the many challenges!

GRATTON's avatar

This may be your most complex and interesting study yet. Congratulations.

FilmFOIA's avatar

One thing I’ve found in the UCLA reports is the use of a flat “median” number which is trumpeted in some of these headlines (e.g. 2022) but is obviously conceptually unhelpful because it means comparing tiny samples with radically different composition in types of film. Looking at “median global box office”/mean number of locations for 2025 when the “n” in question is 1 (native american and MENA respectively), 4 (Asian), 3 (Latinx/hispanic), and 82 (white) are self-evidently not going to be controlling for the relevant variables. There’s nothing meaningful about saying “the median WWBO /mean INT market count of Ballerina, Fantastic Four and Phoenician Scheme (my read on 2025 hispanic lead films) compared to median/mean of Sinners, One of Them Days, Cap 4, Him, Woman in the Yard, and the smurfs reboot” (African American leads). The desire for yearly data just doesn’t match the usefulness of the data. You need to either do more complicated matching or expand your dataset to let year to year data quirks smooth out.

Hollywood Gadfly's avatar

I think something I found difficult in my critique of UCLA was how challenging and assumption driven even the racial categorization process was. One concrete example: they included Black Latinos, which they categprozed as Latino, but not Black. Unfortunately we can’t take even those basic data collection numbers for granted in studies that have such a strong activist bent.

Hollywood Gadfly's avatar

That makes a lot of sense. Looking into these studies I also found it a bit strange that they’re published annually, which incentivizes them to treat noisy data like a giant headline.

FilmFOIA's avatar

to be fair, there is at minimum a nuts and bolts logic to it - you're doing a decent sized data coding lift each year with a number of interns who are presumably contracted for the year. there's just a tension between publishing the descriptive data yearly and running the data analysis (which, especially given inconsistent lag time between production and distribution, really should involve a lot more rolling or historical data). I think USC's "deep dive" into a new topic each year (thus being able to utilize their long term data) is a good way to mitigate these problems.

Hollywood Gadfly's avatar

Fair point. It still doesn’t justify a lot of the conclusions they drew. I looked into the Annenberg Inclusion In The Director’s Chair study as well, which I found a lot of issues in. Which deep dive do you mean? Are there others you find good?

FilmFOIA's avatar

by this I meant the various more in depth reports on a specific topic they'd publish alongside the core one. Honestly, my general approach is to just ignore the intro/conclusions sections and just go directly to the raw data which is just content no one else was creating (including lists of all female or black directors clearing the minimum box office filter). I give them a lot of credit for being willing to publish null/contradictory results as mentioned by Stephen above because it comes alongside this very obvious advocacy lens being applied to the data.

Your Sundance counterpoint/critique is a good one but I think it's pretty rare than I encounter data they generate that I'd consider GIGO. I think the usefulness of basic data collection can be underrated in general. It's really useful to have a source you can trust say "here's the notable age/gender aggregate casting dynamics over time" or "Here's a list of Women Directors 2007-20XX" clearing a pretty modest theatrical release threshold.

If anything, I might push conclusions in the opposite direction - if the state is paying a few million dollars to help generate reports with high quality descriptive data, I'd love for a way to get access to that data instead of having it inherently mediated through the summary reports of the data creator. A study's official report and press release will always lead headlines but raw data also matters. If you think a summary table is being generated by a "n" that's too small to be meaningful, it really does matter if you can generate an alternative table that attempts to slice the data differently or if you're left with the press release v. starting from square 1 of data collection.

This is also why I'm excited by the type of point Stephen Follows makes in the coda of the piece. Especially if you're willing to concede that you're only doing an initial bit of exploratory data analysis, AI really is a gamechanger in lowering the barrier to entry. e.g. in that post I made above, it would have taken me a lot more time to figure out that the Smurfs reboot was coded as having an AA lead in the UCLA database if I couldn't bounce ideas off of the AI after my initial attempt at hand coding didn't quite replicate the numbers they produced. That's obviously very different from having a reliable full dataset but even the lazy version I did there is useful as a stress test.

Hollywood Gadfly's avatar

Great points! Not publishing the datasets is a big issue I’ve found with several of these studies, which force me to make educated guesses as to their assumptions.

I agree the data collection is valuable when done well, I just wish they would stop leaping to such big conclusions based off of it.

I would love to have access to the raw data as well, it’s a public good!

Patrick Quinn's avatar

Very intriguing results.

Hollywood Gadfly's avatar

Wow, amazing work! I hope this leads to higher quality research and more accountability!

Andrew Truong's avatar

above all else, this is an excellent primer on statistical literacy.

Vanessa Hope's avatar

Put this together with Christopher Nolan’s white male dominated films with dead wives, waiting wives or girlfriends with minor, barely developed characters and his “creative” choices suddenly look a lot more big budget studio driven economic. IMHO, we need him to stop dominating our industry with his patriarchal militaristic storytelling. It is bad for culture and bad for progress.

An Actor Explains's avatar

I was a professional actor for over 20 years: I saw people from every culture, race and station get hired for everything from radio to computer game voice over!

- Just look at 80's commercials: they're stock full of female protagonists, African Americans, Asian men and women, people of all backgrounds, like Italians and Japanese in bubble gum commercials, beer ads and food adverts. -

From what I've lived: inclusion was never an issue. It only became one when Activism started to be pivotal for our industry's image. And now, we don't serve Entertainment. Now we serve partisan political agenda...