To analyze President Donald Trump’s spoken words over the last seven months, PolitiFact used a combination of artificial intelligence analysis, standard software and human review. Transcripts were sourced from an authoritative database, and all AI work was reviewed by humans.
Our process worked like this: We first gathered presidential transcript records from PolicyNote, a legislative tracking platform. The database collects all of Trump’s spoken, recorded words and written interviews, which are first transcribed by AI and then proofread by PolicyNote employees.
We used software to separate Trump’s words from other speakers, such as Cabinet members and journalists, to preserve the context. We searched the database of transcripts between Jan. 1 and July 31 to identify times when Trump talked about construction projects such as the White House ballroom, Reflecting Pool or general beautification.
Next, we created a list of 82 topics to categorize Trump’s spoken remarks, including the construction projects and others such as the Iran war and the economy.
We used OpenAI’s Codex, an AI agent, which received one transcript at a time, the topic definitions, and a specialized set of instructions on how to categorize the text by topic. Codex assigned topics sentence by sentence, using the context of the surrounding remarks. It also split sentences if Trump changed topic mid-sentence. Codex could not create new topic labels, but could suggest new labels for human review if none fit. After this process, the AI-generated labels were reviewed by PolitiFact journalists, who could make changes to the topics or reject suggestions.
A separate process classified Trump’s comments as responses to questions or statements he made without being prompted. It used PolicyNote’s speaker labels to identify questions from reporters and interviewers, then treated Trump’s remarks that followed as responses.
The word count for each topic was saved to a database with human review that ensured no missing text or overlapping labels and that the calculated totals used for our analysis matched the transcripts.
— Caleb McCullough and Loreben Tuquero