whoashish115

Ashish Kumar

AI Systems & Research Engineer, Full Stack Developer

Undergraduate Student, Class of 2029

I research and build AI systems end-to-end, from neural network training, world models, generative modeling, reasoning, and multimodal intelligence to GPU kernels, distributed systems, ML infrastructure, and inference.

Email : hello@whoashish.com
Ashish Kumar profile photo

I wrote my first "Hello, World!" when I was 11, and that curiosity eventually grew into a deeper interest in ML, deep learning, and AI. Since then, I've explored LLMs, diffusion models, multimodal learning, NLP, CV, RL, world models, TTS, generative modeling, and agents, with a strong interest in understanding how different kinds of models learn, represent information, and generate outputs.

I especially enjoy making models from scratch. Whether it's an LLM, diffusion model, NLP system, TTS model, or another deep learning model, I like starting from the idea, implementing the architecture, preparing the data, training it, and seeing how far I can take it. I experiment with different architectures, objectives, datasets, and training setups, reproduce ideas from papers, and then modify them to see what else is possible. I also enjoy robotics, AR/VR, and embodied AI, especially where ML can turn these technologies into systems that can understand, generate, and interact with the world.

I believe AGI is possible, and a large part of why I study AI is to understand what it would take to get there. I want to work on the models, algorithms, and systems that move us toward more capable and general forms of intelligence, while exploring how those systems can eventually connect with robotics and spatial computing. For me, the goal is not simply to use existing AI, but to understand it deeply enough to build new models, test new ideas, and contribute to what comes next.

I also play Chess.comChess and do competitive programming, you can watch my profile here on CListCList.

  1. Moonfrost AI —

    A 777M-parameter decoder-only MoE language model with 161M active parameters per token. It combines Multi-head Latent Attention with DeepSeekMoE routing: three of 32 routed experts and one shared expert per layer. Pretraining used about 6B FineWeb-Edu tokens; supervised fine-tuning used 2.7B instruction tokens. Total training time is reported at about 12 H100-hours. The repository includes the tokenizer, training pipeline, evaluation harness, and local streaming inference server.

  2. HyperTracer —

    A C++ path tracer with CPU and CUDA backends, developed from Ray Tracing in One Weekend. It implements physically based rendering and progressive sampling across six procedural scenes. A real-time viewer supports camera movement and scene selection; movement resets accumulated samples and lowers preview resolution to keep interaction responsive. The CPU renderer builds without CUDA, while the CUDA backend supports GPU rendering. The repository includes CMake setup, source, scripts, render commands, and sample outputs.

  3. Wintergreen —

    A C++23 nearest-neighbor search library implementing exact Flat search, approximate IVF and HNSW indexes, Product Quantization, and K-means. Flat search supplies the reference for recall and latency comparisons. The dependency-free core has optional pybind11 bindings for zero-copy NumPy access and releases the GIL during native search. Its CLIP photo-search demo submits the same text query to six indexes to compare retrieval speed and accuracy.

1 post