Human writing mixes short sentences with long ones. Machine drafts tend toward a uniform middle length. Higher variation reads more naturally.
Sentence lengths: 20, 8, 12, 14
Long average sentences make prose heavy regardless of who wrote it. Around 15-20 words reads comfortably; past 25 the reader starts working.
A small set of words appears far more often in generated text than in human writing. Their density is the single most visible tell.
delve · leveraging · robust · seamless · crucial · myriad · facilitate · landscape · underscores · embark · holistic · comprehensive
Phrases like "it is important to note that" add words without adding meaning. Cutting them shortens the text and sharpens it.
it is important to note · In today's fast-paced world
Models punctuate with em dashes several times more often than people typing on a keyboard, where the character is awkward to reach.
Sentences that begin Moreover, Furthermore, Additionally are connective throat-clearing. A few are fine; a pattern of them is a tell.
Moreover, leveraging robust frameworks is crucial for success. · Furthermore, this seamless approach will facilitate numerous outcomes … · Additionally, organisations should embark on a holistic transformation…

該工具測量使散文被閱讀為機器生成的寫作模式,並將每篇散文及其背後的證據報告。它不會輸出模型編寫文字的百分比機會,因為該數字無法誠實生成,而且生成它的危害是真實的。
這就是人工智慧探測器的問題。它們是統計分類器,從表面特徵猜測作者身份,而且它們不夠可靠,以至於 OpenAI 撤回了自己的分類器,因為準確性較低。它們的誤報最容易落在用第二語言寫作的人身上,他們的散文明顯更規則,因此對於分類器來說看起來更像機器。看起來自信的「87% 人工智慧生成」被用來指責學生或拒絕自由工作者,而其背後沒有任何東西可以承載這種分量。任何向您顯示如此數字的工具都會邀請您對拋硬幣做出相應的決定。
可以誠實測量的是文字本身。句子長度的變化,通常稱為突發,是最強的單一訊號:人類寫作將三個單字的句子與三十單字的句子混合在一起,而機器草稿則穩定在統一的中間長度。詞彙密度計算文本達到模型過度使用的一小組單字的頻率。對沖密度計算諸如"值得注意的是"之類的短語,這些短語增加了長度而沒有意義。 Em 破折號率計算字元模型產生的頻率比打字的人高出數倍。連接開機者計算以「此外」或「此外」開頭的句子。
其中每一個都是關於寫作的事實,而不是關於作者的主張,並且每個都報告有產生它的匹配項,因此您可以檢查測量結果而不是信任它。這使得結果對於它實際上有益的東西很有用:編輯。低變化和沈重的對沖會使散文變得更糟,無論是誰寫的,兩者都是可以修復的。
用它來診斷您自己的草稿,查看文件的哪些部分讀作公式化,並了解模式是什麼。不要使用它,也不要讓任何人使用它來決定一個人是否說出了自己工作的真相。一切都在您的瀏覽器中運行;沒有任何內容被上傳、記錄或儲存。
任何散文 - 草稿、文章、提交內容。它會保留在您的瀏覽器中並且永遠不會上傳。
六種測量值,每種測量值都有其值、簡單的英語解釋以及產生它的匹配項。
每個訊號都列出了它實際匹配的內容。如果計數看起來錯誤,證據會告訴您原因。
低變化和沈重的對沖會讓寫它的人寫得更糟。將文字發送給人性化的人,或手工編輯。
報告為變異係數,因此它不依賴絕對長度。集合中最強的單一訊號。
計算大約 40 個單字模型持續達到,每 1000 個單字標準化,因此短文本和長文本進行比較。
增加長度但沒有意義的短語,以及以連接清喉嚨開頭的句子,兩者都列出了匹配項。
每個訊號都顯示實際匹配或原始句子長度,因此可以檢查而不是相信測量結果。
Rewrite the habits that make prose read as machine-written - inflated vocabulary, hedge phrases, em dashes, and connective filler.
Detect and remove invisible Unicode characters from AI text - zero-width joiners, variation selectors, bidi controls, and tag characters.
Generate random numbers in any range, with unique-only and decimal options.
Count words, letters, sentences, paragraphs, and estimate reading time
0 comments
No comments yet. Be the first to share your thoughts!