Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Model Library
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.
Browse and deploy state-of-the-art AI models through the DEVUP Gateway.


GPT-6 Astra is OpenAI's most capable model and the one its documentation recommends by default: built for the hardest end-to-end work rather than for any single task. It carries the newest knowledge of anything in OpenAI's catalogue, against a context window slightly beyond a million tokens and a 128,000-token output ceiling. Reasoning runs across four effort levels — and unlike the generation below it, deliberation cannot be switched off, which is a deliberate signal about what this model is for. It accepts text and images, calls functions, and returns structured output. On DEVUP AI it is reachable through the standard OpenAI-compatible endpoint, billed in Algerian Dinar.


DeepSeek V4 Pro 0813 is the official release of V4 Pro, and the point at which the series flagship became a production agent rather than a preview. It keeps the architecture that defines the series — 1.6 trillion total parameters with 49 billion active per token, a one-million-token context window, and a hybrid attention stack built to make that window practical — and rebuilds the agentic behaviour on top of it. Long-horizon software engineering, terminal automation, security workflows, and tool orchestration all improve substantially over the preview, in several cases by a factor of several. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license.


Gemini 3.7 Flash is Google's most capable Flash-tier model, built for coding and agentic work rather than for chat alone. It is natively multimodal in the widest sense available: text, images, video, audio, and PDFs all go in directly, against a one-million-token context window, with no separate extraction step in front of it. Thinking depth is a request-level setting with three levels, letting the same model serve a latency-critical pipeline and a long multi-step agent run. Google reports large gains over the previous Flash generation on issue resolution, production-ready code generation, complex document processing, and real-world business automation. On DEVUP AI it is the natural choice when your input is not plain text and your workflow has more than one step.


DeepSeek V4 Flash 0731 is the official release of V4 Flash, and it is the checkpoint where the series became genuinely agentic. Built on the same efficiency-first Mixture-of-Experts design — 284B total parameters with only 13B active per token and a one-million-token context window — it adds a substantially rebuilt agentic capability that lifts long-horizon coding, terminal automation, and tool-use scores far above the preview, in several cases past the much larger V4 Pro preview despite activating a fraction of its parameters. Reasoning effort is a per-request setting with three levels, and the checkpoint ships with a speculative decoding module for faster generation. Released under the MIT license, it is the model to reach for on DEVUP AI when an agent has to finish a long job, not just start one.


GLM-5.2 carries 744 billion parameters, activates forty, and holds a million tokens of context that Z.ai describe as solid rather than merely available. The architectural reason is IndexShare: instead of computing a separate index for every sparse attention layer, one indexer is reused across every four — cutting per-token compute by 2.9× at a million-token context, which is exactly where that cost would otherwise become prohibitive. Its speculative decoding layer was improved alongside it, raising acceptance length by up to twenty percent. Coding effort is selectable across multiple levels, and the whole thing ships under MIT with no regional restrictions.


Qwen3-TTS-VoiceDesign invents a speaker rather than copying one. Describe a voice in plain language — age, timbre, accent, mood, pacing — and the model synthesises someone who sounds like that, with no reference recording involved. That is the opposite of voice cloning, and it makes the technique usable where no sample exists or where using a real person's voice is not an option. Qwen report it outperforming a leading closed-source voice-design model on instruction-following and expressiveness alike. It covers ten languages plus dialectal profiles, runs at 1.7 billion parameters, and ships under Apache 2.0.


Qwen3.5-397B-A17B activates seventeen billion parameters out of 397, across 512 experts of which eleven work on any token. Its attention is hybrid rather than uniform — Gated DeltaNet layers interleaved with full attention, so linear-cost layers carry the depth and exact attention appears where retrieval demands it. Vision was trained in from the start through early fusion on multimodal tokens rather than attached afterwards, and Qwen report it outperforming their own dedicated vision-language models on reasoning, coding, and agents. Native context runs to 262,144 tokens, it covers 201 languages, and it ships under Apache 2.0.


Qwen3-Max-Thinking is the reasoning variant of Alibaba's trillion-parameter flagship, and its headline result comes with a condition worth reading: 100% on AIME 2025 and HMMT — achieved with a code interpreter attached and parallel test-time compute applied. That number describes a system rather than a model alone. The figure that describes the model is 58.3 on Humanity's Last Exam, a benchmark built so retrieval cannot solve it. It runs a heavy mode that refines reasoning iteratively across inference steps, invokes search, memory, and code execution adaptively mid-conversation, and was trained on 36 trillion tokens across 119 languages.


Qwen3-TTS is a family of open speech models built on a tokenizer Qwen developed themselves — a multi-codebook encoder running at 12 Hz that compresses speech while preserving the parts most systems discard: breath, hesitation, emphasis, and the acoustics of the room. It covers voice cloning from a sample, voice design from a description, style control at fixed timbre, and long-form synthesis. That last capability carries the most convincing number: ten minutes of continuous speech at a 2.36% word error rate in Chinese and 2.81% in English, where most systems drift. Ten languages, dialectal profiles, Apache 2.0.


Qwen3.6-35B-A3B runs 35 billion parameters and activates three, across 256 experts of which nine work on any token. Its layer layout is published exactly: three linear-attention blocks followed by one full-attention block, repeated ten times — so only ten of forty layers pay quadratic cost, which is what makes a 262,144-token window affordable on a model this size. Grouped-query attention with two key-value heads against sixteen query heads cuts the cache by a further eight times. It carries a vision encoder, covers 201 languages, extends past a million tokens with positional scaling, and ships under Apache 2.0.


Nemotron 3 Super interleaves Mamba-2 layers with mixture-of-experts layers and a selection of attention layers — three mechanisms in one stack, each placed where it costs least. Its routing happens in a compressed latent dimension rather than at full width, which NVIDIA describe as improving accuracy per byte, and it was pre-trained at NVFP4 precision rather than quantised afterwards, making it the first in its family trained that way. It carries 120 billion parameters and activates twelve, was pre-trained on over 25 trillion tokens, and ships with its training data and recipe published alongside the weights.


Nemotron 3 Ultra is the frontier model in NVIDIA's open family — 550 billion parameters, 55 active, 262,144 tokens of context, and Mamba-2 interleaved with mixture-of-experts layers so that long inputs cost a recurrent model's price rather than a transformer's. What sets it apart from almost every model at this scale is what NVIDIA published alongside it: the training data, the recipe, a reproducible evaluation path, and the reward model actually used in its reinforcement learning. That last one lets you inspect the standard the model was trained toward, not only the model that resulted.


Kimi K2.6 is a trillion-parameter multimodal agent that activates thirty-two billion per token, built by continuing the training of an existing base on roughly fifteen trillion mixed visual and text tokens. It ships with two reasoning modes — Thinking on by default, Instant for speed — each with its own recommended temperature, and an optional preserve-thinking mode that carries full reasoning across turns and measurably improves coding-agent work. Native INT4 quantisation is applied during post-training rather than after, so the released weights are the evaluated ones. Moonshot also publishes a verifier for checking that a third-party deployment matches their reference behaviour.


Qwen3-Max is Alibaba's trillion-parameter flagship in its direct-answering form — no reasoning trace, no thinking mode, no deliberation budget to size. Pre-trained on 36 trillion tokens, it splits its 262,144-token window into a fixed allocation: 258,048 for input and 32,768 for output, a ratio you cannot trade against. Independent assessment describes it as not the most conversational or creative model available, and among the strongest for factual and technical prompts — which is the trade it was designed to make. Speed, structured output, and dependable tool use over conversational flourish.


Gemma 4 31B is the largest dense model in Google DeepMind's open family, and it is built to read as much as it reasons. A 256,000-token context window takes entire codebases or large sets of images in one prompt, with a vision encoder that accepts variable aspect ratios rather than forcing everything into a square — which is why it handles UI screenshots, charts, and handwriting recognition rather than approximating them. Thinking is configurable rather than fixed, function calling is native, and the model was pre-trained on more than 140 languages with 35 supported out of the box. Apache 2.0, and it fits on a workstation.


Gemma 4 26B A4B is the sparse member of Google DeepMind's open family: 26 billion parameters in total, roughly four active per token. That ratio is what lets it run on a workstation GPU while reasoning like a considerably larger model. Its attention alternates local sliding windows with full global passes — and always ends on a global layer, so the last thing the model does before answering is look at everything. Image input accepts variable aspect ratios and resolutions rather than a fixed square, and reasoning is a first-class channel in the chat template rather than a tag bolted onto the output. Released under Apache 2.0.


DeepSeek V3.2 is the release where sparse attention stopped being an experiment. Built on a 685-billion-parameter Mixture-of-Experts backbone, it pairs DeepSeek Sparse Attention — a trainable mechanism that attends to a selected subset of past tokens rather than all of them — with a reinforcement-learning framework scaled well beyond the previous generation. The combination produced gold-medal performance at the 2025 International Mathematical Olympiad and International Olympiad in Informatics, and DeepSeek published the actual competition submissions for independent verification. A third pillar, a large-scale pipeline that synthesises agentic training tasks, brings reasoning into tool-use rather than treating the two as separate skills.


DeepSeek V4 Flash carries 284 billion parameters and activates 13 of them per token, against a context window of one million. What makes that window usable rather than nominal is a hybrid attention stack — Compressed Sparse Attention paired with Heavily Compressed Attention — designed specifically to break the quadratic cost that stops long context from being economical. Reasoning depth is a per-request choice across three modes, and the span between them is wider than most model generations: competition mathematics roughly doubles from the fastest setting to the deepest. It ships MIT-licensed, pre-trained on over 32 trillion tokens, with a published technical report.


DeepSeek V4 Pro is the flagship of the DeepSeek V4 series: a Mixture-of-Experts model with 1.6 trillion total parameters, 49 billion of which activate per token, and a one-million-token context window. Its hybrid attention stack was designed specifically to make that window practical rather than nominal, cutting per-token compute to roughly a quarter and key-value cache to roughly a tenth of the previous generation at full context. Reasoning effort is a per-request control with three levels, and at its deepest setting the model reaches its strongest results on world knowledge, long-context retrieval, and multi-step agentic work. Released under the MIT license, it is the model to reach for on DEVUP AI when the input is enormous, the question is genuinely hard, or accuracy outweighs everything else.


Hy4 preview is Tencent's 770-billion-parameter mixture-of-experts flagship, activating 49 billion per token, with a separately-counted multi-token-prediction layer built in for speculative decoding. Its training data came from people who ship — software engineers, game developers, finance analysts, and security experts inside Tencent, with material built around the work they actually deliver. And the evaluation matches: 163 internal experts rated outputs on 203 engineering tasks in a blind comparison, with the full win, tie, and loss breakdown published rather than only the margin. It is explicitly an early release, and Tencent say why.


Ming-Image-0.1-Design generates the kind of image most models get wrong: the kind with words in it. UI screens, infographics, posters, and text-rich compositions, at six billion parameters, with RGBA output and genuinely transparent backgrounds rather than white ones. Its sampling configuration is unusual — twelve steps at a guidance scale of 1.0, which means no classifier-free guidance at all, because the model was trained to arrive rather than to be pushed. Two resolution buckets, 1024 and 2048, and MIT licensing with no conditions attached.


Ming-Image-0.1-Design-Layer runs image generation backwards. Every other image model in this catalogue turns a description into a picture; this one takes a finished, flattened design and pulls it apart into separate transparent RGBA layers — the background, the shapes, the text, each on its own canvas and each independently editable. You supply the image and a layer plan saying how many layers you want, and it returns them alongside a recomposed version you can check the result against. Six billion parameters, twelve sampling steps, MIT licensed.


Claude Opus 5.5 is the first model in its family where the default reasoning effort went down rather than up — from high on the previous generation to medium here — because it reaches the same conclusions with fewer tokens. Customers evaluating it report its lowest effort setting matching or beating the previous model's highest. Thinking is adaptive and always on, across a one-million-token window with 128,000 tokens of output, and it reads images as well as text. Anthropic publish its benchmark results with production safeguards enabled, noting that this likely lowers the scores.


GPT-6 Luna is the tier OpenAI describe as handling clerical work — summarising, extracting, answering — and it carries the same 1.05-million-token window, the same reasoning range from none through max, and the same caching improvements as the model above it. Two things distinguish it. Its knowledge cutoff is more recent than the more expensive tier's, which inverts the usual assumption. And its output-token reduction over its predecessor exceeded the fifty percent headline, which matters most on exactly the high-volume work it was built for.