Uncensored Investigative Journalism

The Pangram Paradox: The Problem with AI “Detecting” other AIs

AI-powered “detectors” of AI-authored content have been promoted as a way to protect human writing online in the age of AI slop. Unfortunately, that’s not what’s really happening.

Op-ed: This is an opinion piece published on Unlimited Hangout. Most articles published on UH are not opinion pieces, so we are clearly labeling those that are.

It should be no secret to readers of Unlimited Hangout that I have a very negative view of the impact that Artificial Intelligence (AI) is having on independent media and media in general. As someone who has been impersonated by AI daily on large video platforms for a few years now, I have come to acutely feel that AI is becoming a dangerous force that is making it increasingly difficult for well-meaning readers/viewers to distinguish between what is real and what is false. I don’t mean just for those hoping to view my content, as the personal issue I am facing is a microcosm of a much larger problem.

As a result, one might think that I would have been pleased, if not relieved, to see Substack, a large platform for professional writers, partner with an “AI detection tool.” Not only am I not pleased, but I have now taken the time to write an opinion piece explaining why I think this is bad and to also explain where I think a lot of this is going.

Substack’s new partner and, apparently, the AI detection tool of choice for most major media outlets right now is called Pangram. Pangram was founded in 2023, originally under the name Checkfor.ai, by Max Spero and Bradley Emi. Both Spero and Emi are Stanford graduates who, prior to Pangram, worked for companies like Nuro and Tesla where they both worked on AI self-driving car software (i.e. replacing human drivers with AI software). Both founders “have long been excited about the potential for AI to transform society” and say they created Pangram “to mitigate new issues caused by the proliferation of powerful generative AI models.”

I find it interesting that two Stanford graduates passionate about AI’s potential to “transform society” who then worked to facilitate AI-powered, human-free driving software decided to quickly pivot to “mitigate” issues caused by generative AI. A lot of the issues caused by generative AI were actually predicted years ago by figures like former Google CEO Eric Schmidt (Pangram co-founder Max Spero previously did work for Google). Schmidt, in a 2021 book authored with war criminal Henry Kissinger, essentially predicted that AI’s “transformative potential” would cause what he called “cognitive diminishment” in the masses and make it almost impossible for people to distinguish between what is real and what is not on their own. Perhaps unsurprisingly, Schmidt and Kissinger’s preferred “solution” to this problem was that we become reliant on AI to distinguish for us what is real and what is not –– in other words, to let another AI become the arbiter of our reality.

Pangram co-founders Max Spero (left) and Bradley Emi (right) – Source

I have written before about this Schmidt-Kissinger book and its predictions, and also noted that there these predictions and warnings –– given the power networks in which both men are/were enmeshed –– were more likely to be disclosure of what had already planned with respect to AI rather than warnings about what might happen upon the mass adoption of AI. Nevertheless, the Schmidt-Kissinger “warnings” were unusually prescient. In addition, their proposed solution is very similar to the what Pangram is currently offering –– in that Pangram uses AI to recognize AI vs. human writing. Pangram thus becomes the currently preferred AI arbiter of what is real human writing and what is not.

There has certainly been a successful marketing campaign to convince us that Pangram is the “most accurate” AI writing detection tool on the market. Headlines of articles promoting Pangram essentially communicate that it will definitely tell you whether AI wrote something. If you read past the headline, which many don’t, you will see the fine print that Pangram will also mark something as AI generated even if AI was only used to edit a human-generated piece of writing. Pangram’s AI-generated label doesn’t really distinguish between the two, even though, as co-founder Max Spero acknowledges, that label “changes how people approach the text.”

The fact that Pangram can’t meaningfully distinguish between a text fully generated by AI or one simply edited by it reveals that it is not an “AI detector” at all. As noted by Tim Requarth back in April:

The term AI “detector” is a bit of a misnomer, as tools like Pangram don’t scan a database of known AI text for a match; they make a guess about likely authorship based on differences in patterns between AI-generated and human-generated text. Rather than a detector, thinks of it as an “AI inferencer” or more colloquially, an “AI best guesser.”

As an inference tool, Pangram is vulnerable to everything that makes inference so difficult: small sample sizes (a single article rather than a corpus), small effect sizes (as the statistical distance between human and AI writing shrinks), and confounds (a human who reads and mimicks AI-generated text before writing her own).

Indeed, the way Pangram works is by making inferences about whether a given piece of writing features enough of the words, phrases, sentence construction, etc. that are currently common in AI-generated text. However, many humans writers, long before generative AI, used a lot of those words, phrases, etc., too. As a consequence, we now have writers who are altering their entirely human writing style in order to please the Pangram gods and hope they won’t be falsely labeled as AI slop purveyors.

A good example of this was recently discussed in a piece by Alex Kirshner for Slate. In that article, Kirshner writes:

Whatever the future holds for the accuracy of A.I. detectors, I am unsettled by what both the A.I. writing and the detectors are doing to me right now. I worry about how A.I. writing and efforts to sniff it out are changing fully human writing. I write all week, every week, and can feel this story’s arsonists (LLMs, particularly when used the wrong way) and firefighters (A.I. detectors) weighing on my process in negative ways.

Like many journalists, my writing style comes from the books, news articles, blogs, and (yes) social media posts that have latched on to my brain over the years. A.I. writing is now in my water supply, whether I can identify each new gulp or not. It feels inevitable that I’ve started to internalize its tics, even the most infamous ones, like incessant em-dashing and overreliance on the “It’s not x but y” structure. For one thing, I fear this will make me a worse writer, dragging a talent that was good enough to get me hired here (and elsewhere!) back toward the A.I.-generated median. Scarier yet, it’s possible that someone—or, heaven forbid, some A.I. detector—might mistake my text for ChatGPT’s. I have started to self-police against this outcome. A few times lately, I’ve caught myself modifying my own work to sound “less like A.I.” In the process, I’ve probably picked words I wouldn’t have settled on otherwise. It feels dystopian—and I doubt I’ve succeeded. 

I would go a step further than Kirschner and say that this outcome is definitely dystopian. It becomes more so when one considers that the supposedly “telltale” patterns of AI-generated text change as new, improved models are released, particularly as they are literally ripping through mountains of obscure, human-written books to train them to write more and more like humans. Meanwhile, the actual humans, to keep a job in a world where the pool of writing jobs is already shrinking rapidly, are scrambling to not sound like something that is rapidly being trained to sound like them as much as possible. Perhaps by design, this is not a game that humans are meant to win and, as I will explain later on, that means it is a game that writers should quickly stop playing.

A screengrab from Pangram’s website – Source

Before we get there, there are a couple of other things to note about Pangram and other AI “detection” companies. One is the issue of false positives. Yes, it is true that Pangram is considered by ostensibly independent researchers (and also itself) to have the smallest false positive rate in the industry. For the purpose of this example, let us use Pangram’s claim that only 1 in 10,000 Pangram-generated labels regarding AI or human authorship are false positives, meaning the text is labeled AI generated when it is not. On Substack, which now has Pangram embedded within it, we do not know exactly how many posts on average are made a day, as Substack does not disclose, but we can guess. The site itself claims that over 1 million posts are “discovered by potential subscribers in the app every day.” The site has over 2 million active publications, most of which publish at least weekly. Let us say, for the purpose of this example, that those 2 million publications do publish weekly. Then, 200 of them, per Pangram’s rate, will be marked as false positives every week. Over a year (52 weeks), we get over 10,000 Substack articles that are falsely labeled as AI, even though they are not. While this figure is still certainly smaller than what Pangram’s competitors offer, if we assume that those 10,000 false positives affect different publications, that could be 10,000 human writers being falsely labeled as AI slop, thereby affecting their reputation, revenue and, perhaps most crucially, their trust with their audience over that calendar year. I see it as a push to outsource trust about a writer’s authenticity from being between a writer and their audience to being between Pangram and a writer’s audience.

While these are statistics used by Pangram and its partners for marketing purposes, behind those numbers are real human writers facing an accusation from Pangram’s AI that could ruin their business. As this effect stacks over years, it could theoretically make the pool of human writers whose work has not been labeled as AI by Pangram smaller and smaller as time goes on. It is also fair to say that the false positive rate that Pangram cites is also the lowest one that has surfaced in research about its product, and it is cited most frequently because the company considers it most favorable to its interests. It is entirely plausible that the actual false positive rate is higher. There is also the added issue that if you run the same text through Pangram more than once, you’re likely to get highly variable –– as opposed to consistent –– results, making Pangram’s verdicts look more based on the luck of the draw than any meaningful “detection” of AI use.

Another important angle to explore with respect to Pangram is the potential for this tool to be deliberately wielded as a weapon for the purpose of censorship and/or discrediting a given journalist or outlet. Max Spero, Pangram’s co-founder, has notably been accused of using his Twitter account to call out specific outlets and journalists for using AI, which has –– in recent months –– generated major headlines and controversy. I don’t really think it’s relevant for the purpose of this article to investigate whether the claims levied against those journalists and outlets are true or not. As recently noted, Pangram does generate false positives, but it is also possible that these journalists did use AI in some way but did not disclose it. However, what is relevant here is that Spero in particular has used the verdict of his company’s AI “detector” to make weaponized claims against specific journalists and outlets.

As a long-time writer in independent media who has consistently written about controversial topics before it later became safe to do so (see my early work on Epstein’s Israeli intelligence ties, Peter Thiel and Palantir, among other topics), I am not a fan of mainstream or “legacy” media. On more than one occasion, they have either defamed me or used my research and writing without citing me. A good part of my career has also been spent picking apart the often misleading narratives they promote. As a result, I tend to be skeptical about the authenticity and veracity of the articles published by many of the same large outlets that Spero has targeted. However, there is another Spero quote that may get your attention if you have been closely following independent media over the past ten years.

In an article published by TechCrunch about Pangram’s recent raising of $9 million dollars from investors, Spero discussed wanting to start Pangram to mitigate new issues that emerged following the public release of generative AI tools. As an example of what types of new issues had emerged as a result of AI’s growing use, Spero offered “LLM-powered Russian disinformation campaigns and UAE-influenced campaigns on Twitter.”

Perhaps you remember the time, beginning roughly ten years ago and continuing for many years since, that any piece of reporting from independent media that was deemed inconvenient to the powers that be (particularly about US proxy wars abroad) was falsely labeled as “Russian disinformation.” Perhaps, more recently, you may remember pro-Israel influencers and media outlets similarly claiming that criticism of Israel in recent years has allegedly been driven by funding from Qatar, as opposed to public outrage over Israel’s live-streamed genocide of Palestinians in Gaza. Or what about the newer claim that LLM-powered bot accounts from China that had no reach whatsoever are really secretly behind Americans’ growing distrust of Big Tech and discontent over data centers?

If you have been following long enough, you may have noticed a trend whereby critics of certain power structures are routinely and falsely labeled as state-sponsored propagandists or “bots” as a way to discredit, not only the critics themselves, but also to discredit the broader concern –– be it the impact of data centers or imperial wars. I do not know if Spero’s intent with Pangram is to similarly weaponize his company’s product in a similar way, but it is certainly possible.

However, the fact that Spero’s Pangram is partnered with Newsguard certainly increases that possibility. As I wrote back in 2019, Newsguard is a dubious “fact checker” that was initially created to discredit independent media outlets critical of US foreign policy, while giving admitted US propaganda organs like VOA a pass. It is also worth mentioning Newsguard’s national security connections — for example, it originally counted the former heads of the CIA and DHS under George W. Bush among its advisors on determining what is “fake news” and what is not. One of their top investors is the founder of Kroll Associates, often referred to as the “private CIA,” Jules Kroll. Now, with Pangram, Newsguard is determining what is “AI slop” and what is not and has used Pangram to launch a “real-time ‘AI content farm’ detection” service.

Newsguard’s original Advisory Board – Source

What better way to digitally nuke a group of reporters and any accompanying wave of public sentiment than by using a now “trusted” AI detector to label both the writers and the sentiment as inauthentic and machine-generated. The potential is for Pangram to become a new way for censoring human-authored, yet inconvenient content, particularly if content labeled AI-generated eventually becomes algorithimically demoted over what Pangram labels as human-generated. The public will be told that the “AI slop” is being buried, when — in reality — it is dissent.

The way it looks to me is that Pangram is currently in the phase of normalizing the ubiquity of its product and building “trust” with those who use Substack or any other platform (or browser plug-in) that integrates Pangram. During this period, readers will be trained to see a Pangram label and trust its results with no further information needed. This is how a company like Pangram can raise $14 million in financing (including from one of Anthropic’s main investors) for a product that it is available for free. As has been said before about online products, if the product is free, then you are the product. In this case, I think that training you to outsource your “trust” to Pangram from the person or outlet you are reading is the product. Once you trust an AI “detector” tool, you will trust its verdicts, perhaps even over the writer you had grown to trust.

Pangram’s co-founders did not make Pangram to protect human-generated writing and speech. Pangram’s co-founders have long been excited that AI will “transform society,” and immediately prior to Pangram, were developing algorithms to replace human drivers with AI. As noted above, they created Pangram to stop what they referred to as state-sponsored influence operations among other “issues” that emerged upon the public release of generative AI tools. With their former Big Tech employers having public sentiment move against them, whether on data centers or “smart” glasses or something else, Pangram’s co-founders may offer a new way of doing away with accounts deemed too critical of AI’s “transformation” of society. Maybe not, but forgive me for being skeptical.

To conclude, I would like to again return to the Schmidt-Kissinger book on AI and its predictions. In independent media, many warn of the Hegelian dialectic, also known as “problem-reaction-solution.” That is, the powers that be first create the problem, then wait for the public to react and finally offer the public the policy outcome they always wanted as the solution. It certainly is convenient that Schmidt and Kissinger, two very terrible people who both helped create this mess in different ways, presciently predicted that AI would make it so people could no longer tell what was real and what was not. Now their prediction has become our reality and, right on cue, the company Pangram has stepped in to solve the AI slop problem (from which its own investors profit) by offering us what was also the Schmidt and Kissinger’s preferred “solution” –– people relying on an AI arbiter to tell them what is real and what is not.

As I wrote earlier, the obvious way out of this dystopian game is not to play it. Do not trust Pangram’s AI just because it looks “authoritative” and has had some positive PR. It is not a “detector” as it claims to be and its results for the same text are rarely consistent. If you are a reader, ask the writers you are reading to disclose their use of AI. Trust should be built between writers and their readers directly, without an AI middleman and arbiter of authenticity or “humanness.” At Unlimited Hangout, we never use AI for writing, research or editing of our articles. This is made clear in our AI use policy, which you can read on our “About Us” page upon the imminent launch of the newly redesigned Unlimited Hangout website (In the meantime, you can read it here). We created this policy after helping create the IMA AI Transparency Pledge, which encourages independent journalists and outlets to voluntarily disclose if you choose to use AI and how.

Just like any other human writer, all I can offer to my readers is my research, my sourcing and my authentic analysis and, in the case of this article, my opinions. I am not going to run my articles obsessively through Pangram hoping for a “100% human-generated” result. I am not going to let fear of an unfavorable Pangram result change how I write and make me afraid of using words like “delve” or “profound” because Pangram might punish me. I know I write without AI and I have enough trust with my readers that they know I don’t use it in my writing either. Not unlike the entities that have labeled inconvenient reporters as Russian/Chinese “bots,” Pangram assumes that you can’t determine what is authentic and inauthentic on your own and that you need them to tell you.

While critical thinking may be in low supply in the age of AI, don’t let it be. You don’t need another AI to “help” you determine what does or doesn’t feel or seem authentic to you. You don’t need it to tell you what you should or shouldn’t read and who you should or shouldn’t trust. It’s never a good idea to outsource your critical thinking and your opinions to others, especially a non-human AI.

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts
Read More

10 Facts From the UK Government Pfizer Vaccine Guidance that Promote “Vaccine Hesitancy”

Official government guidance has been released in the United Kingdom to assist healthcare professionals in administering the Pfizer/BioNTech vaccine BNT162b2. While the UK government goes to war against supposed misinformation, the official narrative is clearly based on very little to no supporting data from incomplete clinical trials. This article examines the document "Reg 174 Information for UK Healthcare Professionals" and narratives being pushed in the mainstream media that directly contradict that document.