OpenAI is facing renewed scrutiny over the data behind its growing list of mathematical achievements, after a second mathematician publicly challenged the company about the origins of its training material.
The latest criticism arrived just days after a heated dispute broke out over whether the company’s models had benefited from unpublished work. A mathematician has now come forward accusing the AI firm of unethical and “dishonest” behaviour, alongside a lack of transparency about where its training data comes from.
Concerns Raised on Mastodon
In a series of posts on the social platform Mastodon, mathematician Andreas Thom raised concerns that interactions he and his colleagues had with the ChatGPT chatbot, prior to OpenAI’s announcement of its progress, may have contributed to the company’s success in the field.
Thom’s intervention adds weight to a wider debate about the sources feeding advanced AI systems and the extent to which researchers’ contributions are acknowledged. His comments follow closely on the heels of an earlier confrontation over whether OpenAI’s models had drawn on work that had not yet been made public.
Questions Over Transparency
The dispute centres on how OpenAI gathers the material used to train its models and whether that process gives due credit to the mathematicians whose exchanges with the chatbot may have shaped later results. The mathematicians involved are seeking clearer answers about the provenance of the data underpinning the company’s mathematical claims.
The accusations point to a lack of transparency about the origins of the material used to develop OpenAI‘s systems, with the researchers calling for proof that their contributions were not incorporated without acknowledgement. The company’s recent announcements have highlighted its expanding capabilities in mathematics, prompting closer examination of how those advances were achieved.
Two mathematicians have now publicly questioned the company within a matter of days, marking a growing pushback from within the academic community over the relationship between researchers’ work and the outputs produced by commercial AI models.
Source
Image: theverge.com