Inco AI
Inference startup targeting agentic workloads. Serves open models like DeepSeek V4.1 Flash and GLM 5.3 Flash via endpoints that rank first on Artificial Analysis output-speed leaderboards. Their proprietary technique, DFlash 2, is a parallel block-diffusion speculative decoding method delivering over 20% more tokens per pass with roughly 1% added latency. Also ships Splash, an open-source local inference engine optimized for Apple silicon, running Qwen3 models at up to 3x Ollama's decode speed.