Hi, I'm Andrew.
I'm a data scientist at CleanChoice Energy, in Washington, D.C.
This is a garden, not a blog. Things here get planted small and edited later, so a note's date says less than its state: a seedling is a stub or a pointer, growing means there's real substance but I'm still moving it around, and evergreen means it's finished and I still stand behind it.
Most of what's here is work: statistics, machine learning, energy, and a few things about music. Start in the garden, or read about me.
Recently tended
- Hashtag Community Detection on Social Networks
A paper on grouping hashtags into communities using metadata-derived entropy scores, applied to Twitter and Parler.
evergreen Jul 2022 networks, research, trss
- Anomix: Mixture Models for Anomaly Detection
An open source Python package for univariate anomaly detection with mixture models. Narrow, but it does its few jobs well.
evergreen May 2022 open-source, statistics, python, trss
- xView3 Challenge and 7th Place Submission
Leading the TRSS team to 7th place in the Defense Innovation Unit's xView3 challenge — finding dark vessels in synthetic aperture radar.
growing Mar 2022 computer-vision, remote-sensing, trss
- Pitchfork Genre Classification
Can you predict an album's genre from the prose of its Pitchfork review? Mostly, yes — and the failures are the interesting part.
seedling Feb 2022 music, classification, pitchfork, text-analysis
- Pitchfork Data Release and Preliminary Analysis
A scraped corpus of Pitchfork reviews and scores, released publicly, plus a first pass of cleaning and exploratory analysis.
seedling Jan 2022 music, data-release, pitchfork