Published 2 min read
By Brooke Coupal
Topics: Research

Have you ever asked a generative artificial intelligence (AI) tool like ChatGPT or Google Gemini to interpret an image, only to be disappointed by the results? Sometimes a small change to an image is enough to make the AI tool misunderstand what it's seeing. In other cases, an intentionally manipulated image can confuse the tool, causing it, for instance, to identify a banana as an apple.

Ming Shao, an associate professor in the Miner School of Computer and Information Sciences, aims to improve AI tools that learn from and process multiple forms of data, such as text, images, audio and video.

“We want to make AI models more robust and safe,” says Shao, who adds that this will lead to more accurate AI-generated content.

Shao is accomplishing his goal through research funded by a National Science Foundation (NSF) Faculty Early Career Development (CAREER) grant totaling nearly $500,000. CAREER grants are the NSF’s most prestigious awards in support of junior faculty who have the potential to serve as academic role models in research and education.

“I want to extend my knowledge and discoveries from this award to the students in the new Applied Artificial Intelligence and Data Science program,” says Shao, who is overseeing the new Bachelor of Science degree offered by the Miner School. “What we learn from our research can be naturally transferred into the program.”

The Importance of Training Data

Today’s AI models learn from massive amounts of data to generate written, visual and audio content.

“Data is the driving force behind AI models,” Shao says. “But that training data can be contaminated.”

For instance, AI may process outdated text or grainy photos, which could lead to inaccurate output. 

Cyberattackers may also manipulate data to trick AI models – for example, by inputting altered photos. Shao recently published a paper in the scientific journal “Neural Networks” that analyzed attacks on vision-language models – AI systems that process visual and written data to generate responses. The paper was produced as a result of his NSF CAREER-funded research.

“First, we explore different types of attacks, and then we try to make the AI models more robust against such attacks,” he says.

Shao plans to strengthen AI models by training them on the types of attacks they may encounter, making them more resistant to such attacks. He is also exploring whether data in one form, such as written text, could help a model remain reliable when another form of data, such as a photo, is maliciously altered.

To help keep AI systems relevant, Shao is developing a continuous learning model that will enable AI to learn from new data, whether in written, visual or audio form. His model will also help mitigate catastrophic forgetting, which is when AI forgets previously learned information after training on new data.

“You always need to improve your models,” he says. “My overall goal is to make AI models robust regardless of the data form.”

Shao has a team of undergraduate and graduate students assisting him with his research. Two of his Ph.D. students are funded by his NSF CAREER award.

“The students are fascinated by new AI techniques, so we’re constantly learning from each other,” he says.