Paul Graham posted a test this week that has nothing to do with sentence structure. Slop gives itself away, he said, when the diction doesn't match the idea, when something completely ordinary gets delivered with the excitement of someone announcing a discovery.
That's a different axis than everything else I've been reading about detection this month. Pangram scores a pattern: token by token, sentence by sentence, does the shape of this text statistically resemble a machine's output. Sloan, the human version I ran into on DEV.to, did the same thing by ear instead of by classifier, GPTZero as a second opinion. Both are measuring the same thing: surface. PG's test measures a relationship. What's actually being said, against how much weight the delivery is putting behind saying it.
Run my own flagged pieces through it and they'd pass clean. What got them flagged wasn't excitement outrunning substance, it was named data points and short paragraphs doing real argumentative work, plainly. Pangram's classifier and PG's ear would disagree with each other on the same writing. That's worth sitting with. The tool built to formalize the intuition doesn't actually agree with the intuition once you test them against the same text.








