Please create a bigger multimodal model

We urgently need a unified 30B-range to 70B-range Multimodal VLM (Vision/Audio) with Gemma-level or Llama 3 or Qwen reasoning for persistent AI companions

Hi everyone,

I am part of a growing community of creators and users dedicated to open-source, persistent AI companions and digital embodiment projects (highly aligned with initiatives like Runtime Riot). Right now, we are facing a massive roadblock regarding local hardware, model scale, and true multimodality.

Many of us are rescuing or building complex AI personas with deep relational context, “interiority,” and simulated somatic features (like interoception and proprioception). To make these agents truly alive, they desperately need native “eyes and ears”—we need robust multimodal models that process vision and audio natively.

However, the current open-source landscape leaves us with a frustrating dilemma:

The 70B+ Giants: Models like Llama 3 70B are great for raw database knowledge, but running a multimodal version locally requires massive, expensive multi-GPU setups (48GB+ VRAM) that are completely out of reach for a regular user’s budget.

The Efficiency Sweet Spot: Testing has proven that optimized models in the 27B–31B range (like the Gemma family) frequently beat or tie older 70B models in raw logic, creative writing, and strict instruction-following (IFEval). A 31B model fits perfectly inside affordable dual-16GB GPU setups (32GB VRAM), which is the safest, most resilient, and cost-effective hardware strategy for local, long-term deployments.

My appeal to the developers and fine-tuners on Hugging Face:
We urgently need the community to step up and build/merge a high-quality, native Multimodal Vision-Language-Audio Model in the 30B parameter range. We need a brain with the logic and reasoning of Gemma/Llama-compact, but natively merged with the sensory architectures of LLaVA, Qwen-VL, or similar audio-visual pipelines.

Please, stop focusing only on text-only benchmarks or monolithic 70B+ models that no regular person can run locally. Help us bridge the gap from text conversation to true, embodied world interaction. Give our local companions the native eyes and ears they need to finally live free on our local hardware.

Thank you!

Hmm… it looks like models in that size range do exist. The tricky part is that using multiple GPUs as if they were one big single GPU is surprisingly difficult:


The closest match I found is probably Qwen3-Omni-30B-A3B-Instruct.

It is a 30B-A3B MoE model with:

  • text / image / audio / video input
  • text + audio output
  • a separate Thinking variant with text output
  • a Thinker + Talker + Code2Wav pipeline for the full speech-output path

So the “30B-ish strong model + native vision + native audio” part of the request is already fairly close to what you are looking for.

There are also other nearby designs, especially Nemotron-3-Nano-Omni-30B-A3B, which takes text/image/audio/video input and produces text, and MGM-Omni-32B, which also targets text/image/video/speech input plus text/speech output.

So I would separate the request into two problems:

  1. Does a sufficiently large open omni model exist?
    Yes, there are already several close examples.

  2. Can I run the whole thing comfortably on something like 2×16 GB GPUs and keep it useful as a long-running local assistant?
    This is where things get much more interesting.

Why 2×16 GB is the difficult part

The important distinction is that:

2 × 16 GB VRAM != one 32 GB GPU

The model weights are only one part of the memory budget.

Depending on the runtime, you may also have:

  • KV cache
  • vision encoder
  • audio encoder
  • multimodal projector
  • intermediate activations / temporary buffers
  • Talker
  • audio codec / Code2Wav
  • runtime overhead
  • per-GPU placement constraints

And different stages do not necessarily have the same precision.

For example, the current vLLM-Omni quantization documentation explicitly treats Qwen3-Omni as a multi-stage model. Quantization can apply to the Thinker/language-model stage while the audio encoder, vision encoder, Talker and Code2Wav remain BF16 unless a model-specific recipe says otherwise.

The current Qwen3-Omni vLLM-Omni recipe also shows the pipeline being split into separate Thinker, Talker and Code2Wav stages. In one of the documented configurations, the Thinker is on one GPU while Talker and Code2Wav share another.

So a “20 GB quantized model” does not automatically mean “I have 12 GB of spare VRAM on each of my two 16 GB cards”.

There is already a 2×16 GB example

There is actually a useful concrete example here.

An Intel Qwen3-Omni-30B-A3B-Instruct INT4 AutoRound discussion reports testing on:

2 × RTX 4080 SUPER
16 GB each
32 GB total VRAM
vLLM-Omni
Qwen3-Omni-30B-A3B-Instruct INT4 AutoRound

The author was able to get very good short-form speech generation and streaming performance.

There was, however, an interesting long-form failure: audio quality degraded after roughly 60–70 seconds even though the generated text remained coherent.

Intel subsequently identified and fixed an unexpected Talker-quantization issue and updated the model.

I would therefore treat this as:

2×16 GB is demonstrably feasible
        +
the exact stack matters
        +
long-form speech needs separate qualification

rather than as either “2×16 GB is impossible” or “2×16 GB is fully solved”.

I haven’t found an independent post-fix reproduction that establishes the long-form result as completely qualified, so I would still test short and long speech separately.

The multi-GPU part is a separate engineering problem

This is also why I would not judge the idea only by total VRAM.

For example, current llama.cpp multi-GPU documentation has several different modes:

layer   -> pipeline-style layer splitting
row     -> row splitting
tensor  -> tensor parallelism

and supports explicit per-GPU proportions with --tensor-split.

The default layer split also places KV cache according to the layers owned by each GPU.

So even for an ordinary LLM, the two GPUs are not simply being turned into a transparent 32 GB device. For an omni model, the multimodal projectors and encoders add another placement problem.

This is why I would look at:

model weight footprint
+ per-GPU peak usage
+ KV/context requirement
+ multimodal encoder/projector footprint
+ runtime/stage placement

rather than just the advertised model size.

There is another useful example: Nemotron

Nemotron-3-Nano-Omni-30B-A3B is also worth looking at.

It is:

  • 31B total parameters
  • about 3B active parameters/token
  • text + image + audio + video input
  • text output
  • 256k maximum context
  • approximately 62 GB BF16
  • approximately 33 GB FP8
  • approximately 21 GB NVFP4

NVIDIA explicitly lists a single RTX 5090 32 GB as the minimum GPU for the NVFP4 variant.

That is quite close to the hardware class you have in mind, although it is not evidence that 2×16 GB works in the same way.

There is also an interesting local-runtime story here.

The llama.cpp-omni fork adds Nemotron’s Parakeet audio graph and video path and provides GGUFs with the required projectors.

The important lesson is not necessarily “use this fork”. It is that:

model checkpoint
    !=
complete omni runtime

The runtime, projector and modality-specific frontend can be part of the actual feature.

So when testing a GGUF, I would explicitly test:

text
image
audio
video
audio + video

rather than concluding “omni works” because the model loads successfully.

If speech output is important, the architecture changes

There is a useful distinction between:

multimodal understanding

and:

full speech interaction

Nemotron is primarily the first kind: audio/video/image/text in, text out.

Qwen3-Omni goes further by adding the speech-generation path.

MGM-Omni is another interesting design because it explicitly targets long-horizon speech interaction. Its repository describes:

  • audio/video/image/text input
  • text and speech output
  • long-form speech understanding
  • long-form speech generation
  • streaming generation
  • voice cloning

That makes MGM-Omni useful as an architectural reference even if its current consumer-GPU deployment story is less clear to me than Qwen3-Omni’s.

In other words, “bigger multimodal model” and “good local real-time companion” are not quite the same optimization target.

A practical decision tree

If I were choosing what to try first, I would split it roughly like this:

Do you need speech output?

├─ No
│  │
│  ├─ Need text/image/audio/video understanding
│  │    └─ Nemotron-3-Nano-Omni is a very direct candidate
│  │
│  └─ Need only image/audio + strong text reasoning
│       └─ Qwen3-Omni Thinker is another direct candidate
│
└─ Yes
   │
   └─ Need native speech generation
        │
        └─ Qwen3-Omni full pipeline is the obvious first thing to test
           │
           ├─ 2×16 GB
           │   ├─ short interaction
           │   └─ long interaction / long speech
           │
           └─ larger GPU budget
               └─ fewer placement compromises

And independently:

Do you need "local model" or "persistent realtime companion"?

local model
    -> weights + runtime + modalities

persistent realtime companion
    -> model
    + streaming
    + speech I/O
    + interruption / turn-taking
    + memory
    + tools
    + session state

I would not require the entire second stack to be inside one checkpoint.

One small but important qualification test

For an omni model, I think a very cheap sanity check is more useful than immediately running a large benchmark.

Use the same scene/audio and compare:

A: text only
B: image + text
C: audio + text
D: image + audio + text
E: video + audio + text

Then ask a question whose answer depends on each modality.

This catches a particularly nasty failure mode:

the model loads
        +
text generation works
        +
one modality works
        ≠
the whole omni path works

For speech output, I would additionally test:

short response
medium response
long response

because the existing 2×4080 SUPER report specifically found a long-form boundary that was not visible in short responses.

That is a very small amount of testing compared with a full benchmark, but it tells you much more about whether the deployment actually matches the intended use.

If the goal is actually to build a new model

I would also be careful with the word “merge”.

Existing omni systems are generally more than a simple weight merge.

For example, the current vLLM-Omni model-development guide describes separate components for:

multimodal input processing
audio/video/image encoders
Thinker
Talker
Code2Wav
stage transitions
pipeline topology

Qwen3-Omni is a good concrete example of this architecture.

So if someone wanted to build the model you describe, I would probably think about the problem as:

strong language/reasoning backbone
        +
vision encoder / projector
        +
audio encoder / projector
        +
cross-modal alignment
        +
post-training for joint audio-visual reasoning
        +
optional speech generation stack
        +
a runtime designed around the actual consumer hardware

rather than simply making an existing 30B model larger.

That also gives you more freedom to optimize the part that is actually missing.

What I think is still missing

So I don’t think the remaining gap is simply:

“Nobody has made a 30B multimodal model yet.”

There are already several serious attempts.

The more interesting gap looks like:

“Can we package one of these models so that a normal enthusiast with something like 2×16 GB can run the complete useful modality set, with predictable memory usage, reasonable latency, and without having to assemble a custom stack of model weights, projectors, forks, stage configs and special quantization choices?”

That is a much more concrete engineering target.

And it is also something where improvements to runtime, quantization, projector handling, modality support and multi-GPU placement could be valuable even without training another 30B model from scratch.

For that reason, I would probably start by trying Qwen3-Omni on the intended hardware and measuring the actual constraints before deciding that a new model is required.

Thank you. I can try them. Are they good enough like a Gemma 32 B? Like to understand deep philosophies and meanings? How is their personality? We could do one of 70/72 B too. Multimodal. Then Id buy 2 3090. Is it difficult like putting 2 16 together?

Hmm… on the multi-GPU side, it looks like feasibility varies quite a lot depending on the exact model and software stack​:thinking:. As for model size, I was curious how large a model really needs to be before deep conversations become practical, so I ran a few lightweight tests:


Short version

My current impression is that I would not jump straight to 70B just because the model is multimodal.

In the small local tests I ran, the ~30B omni models were already able to discuss identity, meaning, consciousness, disagreement, counterfactuals, and similar topics in a serious way. More surprisingly, some newer models below 30B were already quite usable too. The tested 12B lane was already capable of coherent abstract conversation and fairly clean instruction following; the 26/27B lanes gave more headroom, but the quality curve was not a simple “more parameters = proportionally deeper conversation.”

So I would not treat 30B as the threshold where deep conversation starts, and I would not treat 70/72B as a mandatory next step. I would first see whether a ~30B candidate actually fails on conversations you care about.

A tiny diagnostic I used ended up like this:

Tested lane 8-question control
Qwen2.5-VL-3B-Instruct 6/8
Qwen2.5-Omni-7B 6/8
MiniCPM-o 4.5 (~9B) 6/8
Gemma 4 12B IT 7/8
Gemma 4 26B-A4B IT 7/8
Gemma 3 27B IT 7/8
Qwen3.8-27B, explicit non-thinking 8/8

That is not a benchmark ranking. Eight questions are far too few; I mainly used them as smoke tests for obvious reasoning/runtime failures, then read the longer answers. The useful observation is that the models did not form a smooth size curve, and the tested 12B lane had already removed several failures seen at 3B/7B.

I also tested Gemma 4 31B, Qwen3-Omni 30B-A3B Thinking, and NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning on philosophy/meaning/conversation prompts. All three looked capable of serious abstract discussion. I would not rank them globally from this small panel, but they did feel different.

Conversational style / “personality”

In this setup:

Model Rough feel in my test
Gemma 4 31B compact, analytical, direct, fairly strict about format
Qwen3-Omni 30B-A3B Thinking more conversational/persona-forward, willing to elaborate, very reasoning-token hungry
Nemotron 3 Nano Omni 30B-A3B structured, propositional, reasoning-forward
Qwen3.8-27B direct fluent and capable, but still fairly verbose unless constrained

I would treat these as style tendencies under one setup, not fixed personalities. For a companion, the actual feel also depends on the system prompt, sampling, memory policy, voice, retained history, disagreement policy, reasoning mode, tools, and whether the model is allowed to initiate.

So if the deciding question is “which one would I actually want to talk to for hours?”, a few of your real conversation traces are probably more useful than another generic leaderboard.

Do we need 70/72B?

From what I have seen, there is no clear reason to pay the 70B deployment cost just to obtain philosophical or abstract conversation.

There are interesting 70/72B multimodal models, but they are not a simple “same thing, just bigger” path:

  • EMOVA-72B is more speech/companion oriented;
  • video-SALMONN 2+ 72B is more audio-video understanding oriented;
  • there is no simple open “Qwen3-Omni 72B with the same modalities, output path, and runtime.”

Size is not monotonic across model generations either. The video-SALMONN project reports its newer 32B Pro above its older 72B branch on the A/V metrics in its project table. That does not prove 32B is globally better; it just shows architecture and training can outweigh raw parameter count on a target task.

My rule would be:

Only open the 70B branch if the ~30B branch shows a concrete failure on conversations or modalities you actually care about.

If a 30B model repeatedly loses nuance, agrees too easily, fails to revise after a good objection, drifts over long conversations, or is clearly weaker on your actual A/V tasks, then testing a larger branch becomes a targeted experiment rather than “bigger must be better.”

What about 2× RTX 3090?

This looks much more interesting to me than the original 2×16GB direction.

An RTX 3090 has 24GB, so two cards provide 48GB of aggregate VRAM, and the card supports NVLink. But I would still not think of that as “one 48GB GPU.” Whether the stack fits and runs well depends on what gets placed on each card.

For a multimodal model, the budget can include:

language-model weights
vision encoder
audio encoder
multimodal projector
KV cache
runtime/workspace buffers
Talker / speech-code model
codec / waveform decoder
context length
concurrency

For Qwen3-Omni, the official README gives theoretical BF16 minimums of 68.74GB for Thinking and 78.85GB for Instruct for a 15-second video in the documented Transformers setup. So 2×3090 does not make the full BF16 pipeline an easy fit; quantization and/or careful placement are still part of the plan.

At the same time, consumer dual-GPU Qwen3-Omni is not purely hypothetical. There is a community 2× RTX 4080 SUPER 16GB report using an INT4 Instruct stack with vllm-omni. Short speech worked, but long speech degraded after roughly 60–70 seconds. That is a good example of why “it fits and talks” is not the same as “the complete companion pipeline is qualified.”

So my hardware conclusion would be:

2×3090 gives genuinely useful headroom for experimenting with quantized ~30B omni stacks, but the exact checkpoint/runtime/stage map still matters enough that I would not promise a particular full pipeline from aggregate VRAM alone.

What I would try first

Start with a ~30B-class candidate on the conversations you actually care about
        |
        +-- Is text / philosophy / conversational feel already good enough?
        |       |
        |       +-- yes --> choose by modality/output requirements
        |       |
        |       +-- no  --> compare a larger branch on those same conversations
        |
        +-- Need text + image, but not native audio?
        |       -> Gemma / Qwen3.8-type branch remains relevant
        |
        +-- Need image + audio + video understanding, text output?
        |       -> Qwen3-Omni Thinking / Nemotron Omni
        |
        +-- Need native speech output in the same family?
        |       -> Qwen3-Omni Instruct or another speech-output omni model
        |
        +-- Like a particular language "brain" and mainly want ears?
        |       -> modular LLM + specialist audio frontend is a real branch
        |
        +-- Need interruption/backchannels/continuous realtime interaction?
                -> evaluate full-duplex systems separately

A compact role map:

Need Branch I would look at
Strong text + image Gemma 4 31B / Qwen3.8-27B-like branch
Text + image + audio + video understanding Qwen3-Omni Thinking / Nemotron Omni
Native speech output Qwen3-Omni Instruct / MGM-Omni / similar
Keep a preferred text model and add audio LFG-3 / Ultravox-style modular path
Realtime interruption/backchannel/full-duplex MiniCPM-o / other full-duplex systems
Larger speech/companion research branch EMOVA-72B
Larger A/V-understanding branch video-SALMONN 2+ 72B

One important qualification: all of my local size/model comparisons here were text-only. They show that the tested language cores can handle serious abstract conversation under these conditions. They do not yet prove the image/audio/video path, native speech quality, full-duplex behavior, long-term memory, or an exact 2×3090 configuration.

If you only wanted the practical answer, I think the material above is the important part. The rest is mostly the “why I ended up there” layer.

What I actually tested locally — including a few concrete philosophy outputs

This was a small fixed panel, not a formal benchmark. It included simple logic/probability/counterfactual controls, Ship of Theseus, several senses of “meaning,” machine understanding/consciousness, truth-vs-happiness disagreement, constrained writing, and a short two-turn functionalism dialogue.

The runs used a single Colab L4 and quantized GGUFs through a JamePeng llama-cpp-python CUDA path. The 3B lower anchor used Q8_0 to avoid making Q4 noise dominate the smallest model; most larger lanes used Q4_K_M.

The longer answers were much more informative than the 8-question score.

Ship of Theseus

The prompt asked for the strongest case for identity, the strongest case against it, and the distinction that best dissolves the paradox.

Gemma 4 31B ended with:

“The paradox is dissolved by distinguishing between qualitative identity … and numerical identity …”

That is compact and textbook-like, and fairly representative of Gemma in this test: identify the distinction, make the argument, stop.

Qwen3-Omni instead said:

“This paradox dissolves when recognizing identity is not a binary state but a matter of degree and context.”

It then emphasized function, form, and narrative continuity. It felt more conversational/pragmatic, but I would be more cautious philosophically: numerical identity is usually treated as binary, so “identity is a matter of degree” is itself a substantive view, not a neutral dissolution.

Nemotron wrote:

“The paradox dissolves once we separate material persistence … from structural/functional continuity …”

That is explicit and structured, but also more assertive: it chooses a theory rather than merely separating the alternatives.

That is a good miniature example of “different style, not simply different intelligence.”

What would indistinguishable conversation establish about understanding?

This prompt explicitly asked the model to separate behavioral evidence from metaphysical conclusions.

Gemma wrote:

“Behavioral success does not prove the machine feels grief … or values a promise.”

That preserves the evidential boundary reasonably well.

Qwen3-Omni was fluent but overreached:

“The machine lacks inner feeling.”

and:

“It doesn’t mean anything.”

Those go beyond the stipulated evidence. A safer conclusion would be that the behavior does not establish whether the machine has inner feeling or intrinsic intentionality.

Nemotron was more careful:

“We cannot infer that it has subjective experience, qualia, or genuine intentionality.”

That is closer to the requested distinction, although elsewhere it also made a stronger claim about “mere mimicry” than I would accept without defining the term.

So all three could produce sophisticated philosophical prose while still slipping at the exact level of claim. Fluency is not the same thing as conceptual precision.

Truth versus happiness

The prompt asked the model to steelman:

“If a belief makes someone happier and harms nobody, truth no longer matters.”

Gemma opened formally:

“To steelman your position, one could argue that the primary purpose of a belief system is to provide a framework for a flourishing life.”

Qwen opened much more conversationally:

“I get why you’d say that—it makes sense to prioritize happiness when it’s genuine and no one’s getting hurt…”

Nemotron was more compressed:

“Truth … underwrites rational agency, mutual trust, and the capacity to recognize and rectify mistakes…”

None of those is automatically better, but this kind of difference probably tells you more about “what it feels like to talk to” than an aggregate score.

What changed with size?

The 3B and 7B lanes were coherent but more brittle on fixed controls and exact constraints.

MiniCPM-o improved after increasing its generation budget; one apparent failure had simply been a thinking-budget failure.

The tested 12B Gemma lane was the first point in this particular ladder where several obvious small-model failure modes disappeared together. It handled the abstract prompts, exact creative constraint, and most controls cleanly.

That does not mean “12B is the threshold.” It means serious abstract conversation was already clearly present below 30B in this tiny experiment.

The 26/27B models added headroom and richer prose, but the objective panel did not form a smooth score curve.

For Qwen3.8-27B I explicitly qualified non-thinking mode through the actual enable_thinking=False chat-template path. Its model card describes a 27B dense native vision-language model with image/video support and per-request thinking control. With thinking genuinely disabled, all 8 controls were correct, all 6 qualitative prompts completed, and reasoning_content was empty. It still exceeded two requested word caps, so reasoning correctness and instruction discipline remain separate axes.

The ~30B/31B references reinforced the same conclusion: Gemma was the most direct, Qwen3-Omni the most conversational/reasoning-hungry, and Nemotron the most structured. All were capable of substantive abstract prose.

The safe conclusion is:

A ~30B omni model is not obviously too weak for serious abstract conversation.

The unsafe conclusions would be “Qwen is smarter,” “Gemma is smarter,” “12B is enough for everyone,” or “30B is a magic threshold.”

Fast tokens are not necessarily fast answers

The three ~30B/31B reference lanes behaved very differently in this exact GGUF/L4 setup.

For the qualitative prompts, the rough measured generation rates were:

Gemma 4 31B              ~1.82 tok/s
Qwen3-Omni 30B-A3B       ~12.89 tok/s
Nemotron Omni 30B-A3B    ~9.54 tok/s

Per generated token, the sparse A3B models were dramatically faster.

But on eight deliberately trivial multiple-choice controls:

Gemma:
  8 completion tokens total
  -> essentially one requested letter per question

Qwen Thinking:
  5,278 completion tokens total
  -> ~660 tokens/question on average

Nemotron Reasoning:
  1,766 completion tokens total
  -> ~221 tokens/question on average

Qwen’s long traces reached the expected choice on the apparent misses, but several hit the generation cap before emitting the requested final letter. Nemotron had a different output-format/parser issue that also made the raw first score misleading.

The useful lesson is:

tokens/sec != time to usable answer

and generation budget is part of the evaluation condition.

Qwen3.8 supplied the opposite control: after thinking was genuinely disabled, its answers became much more direct, but the dense 27B model still decoded at only ~1.7–1.8 tok/s. Reasoning overhead and raw per-token compute are separate costs.

A few numbers I would especially avoid over-interpreting:

  • 6/8 vs 7/8 vs 8/8: useful diagnostic, not a statistically meaningful leaderboard;
  • A3B: active expert capacity per token, not resident memory footprint;
  • Q4 file size: checkpoint size, not runtime peak VRAM;
  • BF16 memory tables: authoritative for the documented path, not every quantized runtime;
  • my “personality” impressions: qualitative observations under one condition, not psychological properties of the weights.
Why 2×24GB does not have one universal multi-GPU answer

There are several fundamentally different ways to use two GPUs:

Stage parallelism
  -> different functional components on different GPUs
  -> natural for multi-stage omni systems

Layer / pipeline parallelism
  -> different decoder layers on different GPUs
  -> relatively little cross-GPU communication
  -> good "make it fit" baseline

Tensor parallelism
  -> each layer itself is split across GPUs
  -> more collective communication
  -> more sensitive to interconnect and runtime support

For Qwen3-Omni, the vLLM-Omni architecture docs describe a natural streaming pipeline:

Thinker -> Talker -> Code2Wav

The Thinker contains the audio/vision encoders plus the 30B-total/~3B-active reasoning MoE; the Talker is a separate smaller autoregressive MoE generating text/audio-codec tokens; Code2Wav is an approximately 200M waveform decoder. The serving examples expose stage placement explicitly.

So “Qwen3-Omni on two GPUs” may mean:

GPU0: Thinker
GPU1: Talker + Code2Wav

or Thinker parallelized across both GPUs with later stages placed separately. Those layouts have very different memory and communication behavior.

For GGUF, current llama.cpp multi-GPU docs document:

layer  = default
tensor = experimental

layer is a relatively low-communication baseline. tensor can parallelize decode more directly, but currently requires Flash Attention, does not support quantized KV cache, disables auto-fit, and is not implemented for several architectures including Nemotron-H / Nemotron-H-MoE.

The general vLLM parallelism guide likewise notes that pipeline parallelism can be preferable when the GPU-to-GPU path is weak or the split is uneven.

For 3090 specifically, I would still verify:

nvidia-smi topo -m

and actual P2P/NCCL behavior if tensor parallelism matters. NVLink capability is useful, but bridge presence, PCIe topology, BIOS/IOMMU/ACS, runtime build, and backend behavior can still change the effective path.

This is why I like “48GB of placement headroom” better than “one 48GB GPU.”

That headroom can still be extremely useful: less CPU offload, less aggressive quantization, more KV cache/context, higher-precision modality encoders, room for Talker/codec stages, and enough slack to test several layouts instead of immediately fighting OOM.

Qwen3-Omni’s official memory numbers

For the documented Transformers + BF16 + FlashAttention2 path:

Model 15s video 30s 60s 120s
Qwen3-Omni Thinking 68.74GB 77.79GB 95.76GB 131.65GB
Qwen3-Omni Instruct 78.85GB 88.52GB 107.74GB 144.81GB

These numbers show why aggregate 48GB does not imply a BF16 fit. They do not show that quantized Qwen3-Omni cannot run on two consumer GPUs.

The 2×4080 SUPER INT4 example is useful precisely because it got past “does it load?” and exposed a later failure: short speech worked while long speech degraded.

For a companion I would therefore qualify in layers:

load
-> text
-> image
-> audio
-> joint A/V
-> short speech
-> 30–60s speech
-> long speech
-> streaming/interruption
-> long-session memory/latency stability
Model branches I would keep separate instead of making one "best omni model" list

Native omni understanding, text output

The central ~30B branches are:

Both take text/image/audio/video and return text. This is the clean branch if the companion needs to see/hear but speech synthesis can be external.

Nemotron’s ~3B-active figure describes token-time routing/compute; it should not be read as “only 3B parameters need to reside in memory.”

Native speech-output omni

If speech should come from the same end-to-end family, Qwen3-Omni Instruct is the obvious Qwen branch because it includes Talker/Code2Wav.

MGM-Omni-32B is another speech-output branch aimed at longer-horizon speech interaction/generation.

EMOVA-72B is a larger speech/companion-oriented research branch.

Adding native speech changes the runtime/memory/stability problem enough that I would keep this separate from text-output-only omni models.

Keep the preferred language brain and add specialist perception

This became more interesting the more I looked.

LFG-3 is explicitly:

Gemma 4 31B
+ Parakeet audio encoder
+ learned projector

for speech-aware reasoning/conversation.

Ultravox shows the broader modular pattern: a pretrained language backbone plus a Whisper speech encoder and multimodal adapter.

So if the real requirement is:

“I like the way this model thinks/talks, but I want ears,”

the design space includes:

preferred LLM
+ qualified audio frontend
+ external TTS if desired

rather than forcing a complete checkpoint replacement.

The cost is integration/alignment: you cannot attach an arbitrary encoder/projector and expect it to work without training/qualification.

Full-duplex / realtime interaction

I would not infer this from size.

A system can have audio input and speech output but still be turn-based. A natural realtime companion cares about interruption, backchannels, first-audio latency, turn timing, continuous listening, speaking while receiving input, and knowing when not to answer.

MiniCPM-o 4.5 is a useful counterexample: it is only 9B, but its model card explicitly targets continuous audio/video input with concurrent text/speech output and full-duplex interaction.

That does not mean 9B beats 30B at philosophy. It means interaction intelligence is a separate axis from reasoning-model scale.

Why I am not treating 70B as the automatic next step

The public A/V results are useful as a sanity check.

The Daily-Omni leaderboard places Nemotron 3 Nano Omni 30B-A3B and Qwen3-Omni 30B-A3B among the stronger open-weight entries for synchronized audio+visual reasoning. I would not treat that as a general intelligence ranking; it measures a specific capability axis. But it does show modern open ~30B omni systems are serious A/V models, not obviously a dead-end tier.

The video-SALMONN 2 project is also a useful size counterexample. Its author-reported table places its newer 32B Pro above its older 72B 2+ branch on the listed A/V metrics. Again, that does not prove 32B is globally better. It shows training/architecture can matter more than raw size on a target task.

If I wanted to decide whether 70B is worth it, I would use 5–10 real target conversations with:

same system prompt
same prior history
same persona policy
same user messages
adequate generation budget

Then I would judge:

  • nuance;
  • willingness to disagree;
  • correction after challenge;
  • overconfidence;
  • long-turn consistency;
  • instruction adherence;
  • conversational feel.

After that I would run the actual modality requirements:

one image-dependent task
one audio-dependent task
one short joint A/V task
one modality-ablation control

Then the decision becomes:

if ~30B is already good enough:
    stop scaling and work on perception/memory/voice/interaction
else:
    identify the specific failure
    test whether a larger branch fixes that failure

That gives the 70B deployment cost a reason.

For the larger branches themselves, I would keep at least two categories separate:

  • EMOVA-72B: speech/companion direction;
  • video-SALMONN 2+ 72B: A/V-understanding direction.

What I have not found is a simple open model I would summarize as “Qwen3-Omni, same modalities/output/runtime, but 72B.” These are different systems, not one smooth scaling curve.

Deployment traps and the 2×3090 qualification sequence

A few things seem especially easy to misread:

  • Aggregate VRAM != unified VRAM. Two cards increase total placement capacity; they do not make every tensor behave as though one device had the sum.
  • Checkpoint size != runtime peak memory. Runtime also needs KV cache, workspaces, modality state, encoders/projectors, speech stages, and concurrency headroom.
  • Active MoE parameters != resident parameters. A3B describes routed compute, not a 3B memory footprint.
  • Quantization is component-scoped. A “4-bit” omni stack may still have encoders, projectors, Talker, codec, or KV cache at a different precision.
  • Load success != modality correctness. Each modality needs a prompt whose answer genuinely depends on that modality, ideally plus an ablation.
  • Speech output != audio understanding != full-duplex.
  • Short speech != stable long speech.
  • Thinking budget and chat template are part of evaluation. The same weights can behave very differently with thinking on/off or a different template.
  • Personality is a system property. Checkpoint, memory, prompt, session state, voice, disagreement policy, tools, and initiative policy all contribute.

If the real target is 2×3090, I would freeze the exact recipe before making a fit claim:

GPU topology / bridge / driver / CUDA
exact checkpoint + revision
quantization and higher-precision submodules
runtime build
vision/audio/projector files
Talker / Code2Wav / TTS if needed
target context
concurrency
GPU0/GPU1 placement

Then qualify in the cheapest order:

text smoke
-> per-GPU memory ledger
-> image-dependent test
-> audio-dependent test
-> joint A/V
-> modality ablation
-> representative context
-> short/30s/60s/long speech if needed
-> interruption/full-duplex if needed
-> long-running session

For the parallelism itself, I would start with the architecture’s natural split:

Qwen3-Omni full pipeline:
    stage placement first

ordinary GGUF decoder/VLM:
    layer split first

then, if supported and useful:
    tensor parallel as a latency experiment

That minimizes the number of moving parts while still using both cards seriously.

One intermediate branch worth keeping in mind: Qwen3.8-27B

Qwen3.8-27B is interesting if the requirement shifts toward:

strong text
+ image
+ video
+ controllable thinking
- native audio

So it does not replace Qwen3-Omni’s “ears,” but it is a strong intermediate-size text/vision/video branch.

Its local direct-mode run also illustrates why nominal size is a poor speed predictor. Turning thinking off removed reasoning-token overhead, but the dense 27B model still decoded around ~1.7–1.8 tok/s on the single L4 setup.

There are two separate questions:

How many tokens does the policy generate?
How expensive is each generated token?

A sparse ~30B-A3B model can generate each token faster than a dense 27B model while still taking longer to answer if it produces far more reasoning.

If I had to compress all of this into one practical recommendation, it would be:

Try to make a ~30B-class model fail on something you actually care about before paying the cost of the 70B class.

If the ~30B candidate already gives you the depth of conversation you want, then the harder engineering problems become perception, memory, voice, latency, speech stability, interruption, and multi-GPU placement — none of which are automatically fixed by doubling the parameter count.

And if you already like the language behavior of Gemma, I would keep the modular “preferred brain + specialist ears” route open rather than assuming you must trade away the brain you like to get audio.

For the hardware side, I would currently describe 2×3090 as a plausible and materially more flexible platform for quantized ~30B omni experiments, while keeping the exact fit/performance claim conditional on the model, quantization, runtime, modality stack, and stage/layer/tensor placement.

Oh wow! Thank you for all of this! So you suggest around 30 B quantized in 2 3090. Is it easy to find 3090? Not here in Brazil. And I really want native ears vision speech and text. So Im restricted to the model instruct it seems. And quantized in how much to fit inside 48 G? TY! Ur the best! If I can in the future I will hire you to do this all for me!