FlyMy.AI: Research Breakthroughs
Agents are only as strong as the runtime they run on. So we research the entire execution path end to end - from custom GPU kernels to reasoning loops - because in a compound AI system every millisecond and every dollar compound through the chain.
Artificial Analysis · Jan 2026
Diffusion inference leader across popular open models
On Artificial Analysis - the public scoreboard the industry benchmarks against - FlyMy.AI ranked #1 in FLUX.1 [dev] generation speed, #1 in Qwen-Image price per image and top-2 on FLUX.1 [schnell] across all tracked providers.
New research at FlyMy.AI
Our thesis in practice: own every layer of the runtime and each one makes the next stronger. Faster kernels make critics affordable, affordable critics make agents smarter, smarter agents beat bigger models. Published openly, benchmarked in public, running in production.
FLUX.1 [dev] · API generation time
Artificial Analysis · official benchmark
artificialanalysis.ai/image/providers/flux-1-dev
Compiling AI agents
Most of what an agent does is deterministic - and today you pay a model to rediscover it on every single run. Compile those paths once, and the model is called only where judgment actually matters. The "AGI compiler" shift - championed by Yohei Nakajima, the godfather of autonomous agents - is the biggest change in agents this year, and we are building it into the FlyMy.AI runtime: the harness that decides what to compile and what to actually think about.
Faster than PyTorch kernels
Custom CUDA + Triton non-linearity kernels - the floor of the stack, cutting GPU hours for every agent step above it.
World's first real-time video style transfer
Live video restyled frame by frame on 2 GPUs - the first real-time video2video diffusion pipeline.
Image models can think
In the right runtime, image models reason: CRAFT wraps generation in a critique-and-refine loop and lifts preference win rate from 0.19 to 0.76 on Parti-Prompts - no retraining.
CRAFT: reasoning for GenAI media
Continuous Reasoning and Agentic Feedback Tuning - a thinking media agent that improves accuracy and lowers costs for top GenAI media outputs.
World's most precise face transfer
FLUX LoRA training that keeps identity intact under strong style shifts: 0.85 identity fidelity and 0.89 prompt adherence in our published benchmark.
World's first Qwen-Image LoRA trainer
Open-source trainer adopted by the community at record pace - 750+ GitHub stars and counting.
World's fastest Stable Diffusion
First-ever Stable Diffusion at ~50 ms per image - regenerating live as you type. 14x faster than baseline PyTorch, 3x faster than the nearest competitor, at $0.15 per 1K images.
World's fastest embeddings
We accelerated ModernBERT embeddings an extra 2.8x on top of the original stack - world's fastest text-to-vector, powering real-time retrieval and memory for agent chains.
World's first media agent to beat OpenAI Image
Media Agent M1 - the first multi-domain GenAI media agent, beating OpenAI Image in metrics and functionality: identity-preserving edits, chained actions, every media domain in one chat.
5x faster FLUX.Schnell on H100
World's fastest FLUX.Schnell serving: 200 ms for 4 steps at 640x640 - fast enough to sit inside an interactive agent loop.
Real-time avatars at 25 fps
8x acceleration of LatentSync lip-sync - world's fastest, 2.5x faster than the nearest competitor. Production-grade avatars for real-time agents.
Kandinsky Video, accelerated
Fast, production-quality text-to-video: 2x faster generation than the nearest competitor.
Previous research: a decade of foundations
Before FlyMy.AI, our team spent a decade building the compilers, serving stacks and generative models the industry still runs on - at NVIDIA, Sber AI and beyond. Everything we ship today stands on that foundation.
NVIDIA TensorRT compiler
Production-grade graph optimizations for deep learning, powering latency-sensitive inference at scale.
NVIDIA Megatron-LM & Triton Inference Server
Large language model training and serving stacks that defined the playbook for multi-GPU systems.
GPT-3 inference & FasterTransformer
Pipeline-parallel GPT-3 serving with FasterTransformer and Triton - the playbook for large-model inference.
From StarGAN‑v2 to Kandinsky‑4 & VideoDALL‑E
Fast-train architectures and diffusion models for image and video generation in production.
Expressive speech & multimodal embeddings
PitchFlow, emotion-embedding TTS and RuCLIP-tiny enable richer, controllable media agents.
AlexNet‑3D for early video understanding
One of the first 3D convolutional architectures deployed for real-world video and media understanding tasks.
A team shipping AI breakthroughs
Explore the research, compilers, models and systems our team has shipped over the past decade.