Questionable methodology, as it rests strictly on LLMs.
And because "an AI said so" is not evidence, I ran the whole corpus through an old-fashioned natural-language-processing pass (VADER) that does plain word counts and rule-based sentiment.
But then the author states the results of that did not really align with the LLMs' or the whole purpose of the assessment:
On the subtler question of which side a story takes, it falls apart in a way that reinforces the choice to use models for the more nuanced reads: the NLP scores an op-ed condemning bigotry as "negative" because of words like assault and erase, the exact mistake the models are better able to avoid. That gap, between counting words and reading framing, is the whole reason I leaned on models for the parts a word-counter can't do.
You read the same words I did - they were able to explain the differences in conclusion being due to the traditional sentiment analysis falsely scoring anti-bigotry as negative
Yeah it's questionable for sure. Though that said, sentiment analysis is one of the things LLMs are really good at, so it might have more credibility than one would think.
They're too variable to trust. They can produce really thorough assessments but other times fail miserably even at things like basic sarcasm.
Yeah, I'd want to see more about the testing methodology before making a definitive judgement, I think. Judged it once? Nah. Judged many times with one model? Mehh probably still not reliable. Judged many times across different models, with varying prompts? Maybe!
Not to worry I'll simply ask chat gpt if this article is biased
LGBTQ+
All forms of queer news and culture. Nonsectarian and non-exclusionary.
See also this community's sister subs Feminism, Neurodivergence, Disability, and POC
Beehaw currently maintains an LGBTQ+ resource wiki, which is up to date as of July 10, 2023.
This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.