Two and a half years in the past, Abhinav Anand stated he barely knew what AI was past listening to about ChatGPT.
As we speak, the 19-year-old Class 12 pupil from Bihar claims he has constructed a 5.82-billion-parameter multimodal AI mannequin after spending almost Rs 11 lakh from his private financial savings and compute grants.
In a put up shared on Reddit, Anand stated he labored independently with no group, buyers or a proper laptop science diploma. He described years of failed experiments earlier than arriving at what he calls “ArcleIntelligence,” a multimodal mannequin which he claims is able to processing textual content, pictures, paperwork, audio and video.
“Each failure taught me one thing actual,” Anand wrote, recalling earlier makes an attempt at constructing a YouTube analytics app, a voice assistant and an offline AI assistant.
Based on Anand, the mannequin helps picture technology at 512×512 decision, 24kHz speech output and a context window of greater than 2 million tokens.
He additionally claimed the system achieved a rating of 93.45 on OmniDocBench V1.5 throughout non-public testing, although these benchmark outcomes haven’t been independently verified.
Earlier than starting work on the multimodal mannequin, Anand stated he had educated a text-to-video system on his laptop computer with no outdoors funding and later revealed it publicly by Lightning AI as a studio template. {The teenager} stated the mission was funded by private financial savings, RunPod compute grants, DigitalOcean credit and GitHub Scholar Pack advantages.
He estimated that GPU compute alone price his household about Rs 64,000, which he described as a major quantity for a middle-class family in Bihar.
“My father is a authorities officer. My mom is a housewife. This can be a middle-class household in Bihar,” Anand wrote.
He added that the mission remains to be in coaching and that he’s searching for about $35,000 in funding to finish the pipeline. Anand stated he plans to launch the mannequin weights on Hugging Face and finally open-source the complete codebase on GitHub.
“The West has OpenAI. The East has DeepSeek. India deserves its personal,” he wrote.
19 year old from Bihar, no team, no investors, no CS degree — spent $11,560 of personal savings building a 5.82B multimodal AI. 93.45 on OmniDocBench V1.5 in private testing. Trying to release it open source.
by u/That-Bookkeeper-8316 in indianstartups
Netizens response
The claims rapidly drew consideration on-line, with reactions starting from admiration to skepticism. Some customers questioned the technical credibility of the mission and requested for extra transparency across the coaching course of and datasets.
“What proprietary information have you ever used? Why a multimodal mannequin? The explanation I say is, for the quantity you might be asking, a small multimodal mannequin shouldn’t be going to be manufacturing use on any process,” a consumer wrote.
“I received’t speak in regards to the mannequin however the pitch. 19yo, younger expertise signalled. Bihar, center class, poor nr sympathy signalled. No CS diploma, arduous employee signalled,” one other wrote.
Others raised doubts in regards to the public proof shared on-line.
“Your Twitter account doesn’t say something about your journey. There’s no put up with actual engagement, and the Lightning AI reach-out you point out is a screenshot which could be simply doctored,” a 3rd wrote.
A number of customers additionally accused the mission of relying closely on AI-generated code and “vibe coding.”
“Simply checked the supply code, it’s purely vibe-coded, nothing new,” one other added.
Nonetheless, some customers praised the ambition behind the hassle, stating that even making an attempt to coach a big AI mannequin independently from Bihar was notable in itself.