llama.cpp Jobs - August 2026
Search by company, role, stack, location, salary signal, source, and work setup.
-
Preferred Networks
LLM Inference Optimization Engineer
Posted 1 month, 2 weeks ago
Preferred Networks is an AI company based in Tokyo working across the stack, from AI chips and computing infrastructure to LLMs and products. Preferred Networks is designing in-house chips (MN-Core series) and training LLMs (PLaMo series). We are actively hir…
Tech stack
Location
Tokyo, Remote in Japan
Work setup
full-time · Tokyo or Remote in Japan. Both roles require relocation to Japan and visa/relocation support is available.
-
NobodyWho
Software Engineer
Posted 8 months, 2 weeks ago
NobodyWho is making developer tools for running small language models in local-first applications. Our core principle is to ship the model weights along with the application, and then do efficient inference locally and offline, on any device. We run fast on L…
Tech stack
Location
Copenhagen, Denmark