Hmm… on the multi-GPU side, it looks like feasibility varies quite a lot depending on the exact model and software stack
. As for model size, I was curious how large a model really needs to be before deep conversations become practical, so I ran a few lightweight tests:
Short version
My current impression is that I would not jump straight to 70B just because the model is multimodal.
In the small local tests I ran, the ~30B omni models were already able to discuss identity, meaning, consciousness, disagreement, counterfactuals, and similar topics in a serious way. More surprisingly, some newer models below 30B were already quite usable too. The tested 12B lane was already capable of coherent abstract conversation and fairly clean instruction following; the 26/27B lanes gave more headroom, but the quality curve was not a simple “more parameters = proportionally deeper conversation.”
So I would not treat 30B as the threshold where deep conversation starts, and I would not treat 70/72B as a mandatory next step. I would first see whether a ~30B candidate actually fails on conversations you care about.
A tiny diagnostic I used ended up like this:
| Tested lane |
8-question control |
| Qwen2.5-VL-3B-Instruct |
6/8 |
| Qwen2.5-Omni-7B |
6/8 |
| MiniCPM-o 4.5 (~9B) |
6/8 |
| Gemma 4 12B IT |
7/8 |
| Gemma 4 26B-A4B IT |
7/8 |
| Gemma 3 27B IT |
7/8 |
| Qwen3.8-27B, explicit non-thinking |
8/8 |
That is not a benchmark ranking. Eight questions are far too few; I mainly used them as smoke tests for obvious reasoning/runtime failures, then read the longer answers. The useful observation is that the models did not form a smooth size curve, and the tested 12B lane had already removed several failures seen at 3B/7B.
I also tested Gemma 4 31B, Qwen3-Omni 30B-A3B Thinking, and NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning on philosophy/meaning/conversation prompts. All three looked capable of serious abstract discussion. I would not rank them globally from this small panel, but they did feel different.
Conversational style / “personality”
In this setup:
| Model |
Rough feel in my test |
| Gemma 4 31B |
compact, analytical, direct, fairly strict about format |
| Qwen3-Omni 30B-A3B Thinking |
more conversational/persona-forward, willing to elaborate, very reasoning-token hungry |
| Nemotron 3 Nano Omni 30B-A3B |
structured, propositional, reasoning-forward |
| Qwen3.8-27B direct |
fluent and capable, but still fairly verbose unless constrained |
I would treat these as style tendencies under one setup, not fixed personalities. For a companion, the actual feel also depends on the system prompt, sampling, memory policy, voice, retained history, disagreement policy, reasoning mode, tools, and whether the model is allowed to initiate.
So if the deciding question is “which one would I actually want to talk to for hours?”, a few of your real conversation traces are probably more useful than another generic leaderboard.
Do we need 70/72B?
From what I have seen, there is no clear reason to pay the 70B deployment cost just to obtain philosophical or abstract conversation.
There are interesting 70/72B multimodal models, but they are not a simple “same thing, just bigger” path:
- EMOVA-72B is more speech/companion oriented;
- video-SALMONN 2+ 72B is more audio-video understanding oriented;
- there is no simple open “Qwen3-Omni 72B with the same modalities, output path, and runtime.”
Size is not monotonic across model generations either. The video-SALMONN project reports its newer 32B Pro above its older 72B branch on the A/V metrics in its project table. That does not prove 32B is globally better; it just shows architecture and training can outweigh raw parameter count on a target task.
My rule would be:
Only open the 70B branch if the ~30B branch shows a concrete failure on conversations or modalities you actually care about.
If a 30B model repeatedly loses nuance, agrees too easily, fails to revise after a good objection, drifts over long conversations, or is clearly weaker on your actual A/V tasks, then testing a larger branch becomes a targeted experiment rather than “bigger must be better.”
What about 2× RTX 3090?
This looks much more interesting to me than the original 2×16GB direction.
An RTX 3090 has 24GB, so two cards provide 48GB of aggregate VRAM, and the card supports NVLink. But I would still not think of that as “one 48GB GPU.” Whether the stack fits and runs well depends on what gets placed on each card.
For a multimodal model, the budget can include:
language-model weights
vision encoder
audio encoder
multimodal projector
KV cache
runtime/workspace buffers
Talker / speech-code model
codec / waveform decoder
context length
concurrency
For Qwen3-Omni, the official README gives theoretical BF16 minimums of 68.74GB for Thinking and 78.85GB for Instruct for a 15-second video in the documented Transformers setup. So 2×3090 does not make the full BF16 pipeline an easy fit; quantization and/or careful placement are still part of the plan.
At the same time, consumer dual-GPU Qwen3-Omni is not purely hypothetical. There is a community 2× RTX 4080 SUPER 16GB report using an INT4 Instruct stack with vllm-omni. Short speech worked, but long speech degraded after roughly 60–70 seconds. That is a good example of why “it fits and talks” is not the same as “the complete companion pipeline is qualified.”
So my hardware conclusion would be:
2×3090 gives genuinely useful headroom for experimenting with quantized ~30B omni stacks, but the exact checkpoint/runtime/stage map still matters enough that I would not promise a particular full pipeline from aggregate VRAM alone.
What I would try first
Start with a ~30B-class candidate on the conversations you actually care about
|
+-- Is text / philosophy / conversational feel already good enough?
| |
| +-- yes --> choose by modality/output requirements
| |
| +-- no --> compare a larger branch on those same conversations
|
+-- Need text + image, but not native audio?
| -> Gemma / Qwen3.8-type branch remains relevant
|
+-- Need image + audio + video understanding, text output?
| -> Qwen3-Omni Thinking / Nemotron Omni
|
+-- Need native speech output in the same family?
| -> Qwen3-Omni Instruct or another speech-output omni model
|
+-- Like a particular language "brain" and mainly want ears?
| -> modular LLM + specialist audio frontend is a real branch
|
+-- Need interruption/backchannels/continuous realtime interaction?
-> evaluate full-duplex systems separately
A compact role map:
| Need |
Branch I would look at |
| Strong text + image |
Gemma 4 31B / Qwen3.8-27B-like branch |
| Text + image + audio + video understanding |
Qwen3-Omni Thinking / Nemotron Omni |
| Native speech output |
Qwen3-Omni Instruct / MGM-Omni / similar |
| Keep a preferred text model and add audio |
LFG-3 / Ultravox-style modular path |
| Realtime interruption/backchannel/full-duplex |
MiniCPM-o / other full-duplex systems |
| Larger speech/companion research branch |
EMOVA-72B |
| Larger A/V-understanding branch |
video-SALMONN 2+ 72B |
One important qualification: all of my local size/model comparisons here were text-only. They show that the tested language cores can handle serious abstract conversation under these conditions. They do not yet prove the image/audio/video path, native speech quality, full-duplex behavior, long-term memory, or an exact 2×3090 configuration.
If you only wanted the practical answer, I think the material above is the important part. The rest is mostly the “why I ended up there” layer.
What I actually tested locally — including a few concrete philosophy outputs
This was a small fixed panel, not a formal benchmark. It included simple logic/probability/counterfactual controls, Ship of Theseus, several senses of “meaning,” machine understanding/consciousness, truth-vs-happiness disagreement, constrained writing, and a short two-turn functionalism dialogue.
The runs used a single Colab L4 and quantized GGUFs through a JamePeng llama-cpp-python CUDA path. The 3B lower anchor used Q8_0 to avoid making Q4 noise dominate the smallest model; most larger lanes used Q4_K_M.
The longer answers were much more informative than the 8-question score.
Ship of Theseus
The prompt asked for the strongest case for identity, the strongest case against it, and the distinction that best dissolves the paradox.
Gemma 4 31B ended with:
“The paradox is dissolved by distinguishing between qualitative identity … and numerical identity …”
That is compact and textbook-like, and fairly representative of Gemma in this test: identify the distinction, make the argument, stop.
Qwen3-Omni instead said:
“This paradox dissolves when recognizing identity is not a binary state but a matter of degree and context.”
It then emphasized function, form, and narrative continuity. It felt more conversational/pragmatic, but I would be more cautious philosophically: numerical identity is usually treated as binary, so “identity is a matter of degree” is itself a substantive view, not a neutral dissolution.
Nemotron wrote:
“The paradox dissolves once we separate material persistence … from structural/functional continuity …”
That is explicit and structured, but also more assertive: it chooses a theory rather than merely separating the alternatives.
That is a good miniature example of “different style, not simply different intelligence.”
What would indistinguishable conversation establish about understanding?
This prompt explicitly asked the model to separate behavioral evidence from metaphysical conclusions.
Gemma wrote:
“Behavioral success does not prove the machine feels grief … or values a promise.”
That preserves the evidential boundary reasonably well.
Qwen3-Omni was fluent but overreached:
“The machine lacks inner feeling.”
and:
“It doesn’t mean anything.”
Those go beyond the stipulated evidence. A safer conclusion would be that the behavior does not establish whether the machine has inner feeling or intrinsic intentionality.
Nemotron was more careful:
“We cannot infer that it has subjective experience, qualia, or genuine intentionality.”
That is closer to the requested distinction, although elsewhere it also made a stronger claim about “mere mimicry” than I would accept without defining the term.
So all three could produce sophisticated philosophical prose while still slipping at the exact level of claim. Fluency is not the same thing as conceptual precision.
Truth versus happiness
The prompt asked the model to steelman:
“If a belief makes someone happier and harms nobody, truth no longer matters.”
Gemma opened formally:
“To steelman your position, one could argue that the primary purpose of a belief system is to provide a framework for a flourishing life.”
Qwen opened much more conversationally:
“I get why you’d say that—it makes sense to prioritize happiness when it’s genuine and no one’s getting hurt…”
Nemotron was more compressed:
“Truth … underwrites rational agency, mutual trust, and the capacity to recognize and rectify mistakes…”
None of those is automatically better, but this kind of difference probably tells you more about “what it feels like to talk to” than an aggregate score.
What changed with size?
The 3B and 7B lanes were coherent but more brittle on fixed controls and exact constraints.
MiniCPM-o improved after increasing its generation budget; one apparent failure had simply been a thinking-budget failure.
The tested 12B Gemma lane was the first point in this particular ladder where several obvious small-model failure modes disappeared together. It handled the abstract prompts, exact creative constraint, and most controls cleanly.
That does not mean “12B is the threshold.” It means serious abstract conversation was already clearly present below 30B in this tiny experiment.
The 26/27B models added headroom and richer prose, but the objective panel did not form a smooth score curve.
For Qwen3.8-27B I explicitly qualified non-thinking mode through the actual enable_thinking=False chat-template path. Its model card describes a 27B dense native vision-language model with image/video support and per-request thinking control. With thinking genuinely disabled, all 8 controls were correct, all 6 qualitative prompts completed, and reasoning_content was empty. It still exceeded two requested word caps, so reasoning correctness and instruction discipline remain separate axes.
The ~30B/31B references reinforced the same conclusion: Gemma was the most direct, Qwen3-Omni the most conversational/reasoning-hungry, and Nemotron the most structured. All were capable of substantive abstract prose.
The safe conclusion is:
A ~30B omni model is not obviously too weak for serious abstract conversation.
The unsafe conclusions would be “Qwen is smarter,” “Gemma is smarter,” “12B is enough for everyone,” or “30B is a magic threshold.”
Fast tokens are not necessarily fast answers
The three ~30B/31B reference lanes behaved very differently in this exact GGUF/L4 setup.
For the qualitative prompts, the rough measured generation rates were:
Gemma 4 31B ~1.82 tok/s
Qwen3-Omni 30B-A3B ~12.89 tok/s
Nemotron Omni 30B-A3B ~9.54 tok/s
Per generated token, the sparse A3B models were dramatically faster.
But on eight deliberately trivial multiple-choice controls:
Gemma:
8 completion tokens total
-> essentially one requested letter per question
Qwen Thinking:
5,278 completion tokens total
-> ~660 tokens/question on average
Nemotron Reasoning:
1,766 completion tokens total
-> ~221 tokens/question on average
Qwen’s long traces reached the expected choice on the apparent misses, but several hit the generation cap before emitting the requested final letter. Nemotron had a different output-format/parser issue that also made the raw first score misleading.
The useful lesson is:
tokens/sec != time to usable answer
and generation budget is part of the evaluation condition.
Qwen3.8 supplied the opposite control: after thinking was genuinely disabled, its answers became much more direct, but the dense 27B model still decoded at only ~1.7–1.8 tok/s. Reasoning overhead and raw per-token compute are separate costs.
A few numbers I would especially avoid over-interpreting:
- 6/8 vs 7/8 vs 8/8: useful diagnostic, not a statistically meaningful leaderboard;
- A3B: active expert capacity per token, not resident memory footprint;
- Q4 file size: checkpoint size, not runtime peak VRAM;
- BF16 memory tables: authoritative for the documented path, not every quantized runtime;
- my “personality” impressions: qualitative observations under one condition, not psychological properties of the weights.
Why 2×24GB does not have one universal multi-GPU answer
There are several fundamentally different ways to use two GPUs:
Stage parallelism
-> different functional components on different GPUs
-> natural for multi-stage omni systems
Layer / pipeline parallelism
-> different decoder layers on different GPUs
-> relatively little cross-GPU communication
-> good "make it fit" baseline
Tensor parallelism
-> each layer itself is split across GPUs
-> more collective communication
-> more sensitive to interconnect and runtime support
For Qwen3-Omni, the vLLM-Omni architecture docs describe a natural streaming pipeline:
Thinker -> Talker -> Code2Wav
The Thinker contains the audio/vision encoders plus the 30B-total/~3B-active reasoning MoE; the Talker is a separate smaller autoregressive MoE generating text/audio-codec tokens; Code2Wav is an approximately 200M waveform decoder. The serving examples expose stage placement explicitly.
So “Qwen3-Omni on two GPUs” may mean:
GPU0: Thinker
GPU1: Talker + Code2Wav
or Thinker parallelized across both GPUs with later stages placed separately. Those layouts have very different memory and communication behavior.
For GGUF, current llama.cpp multi-GPU docs document:
layer = default
tensor = experimental
layer is a relatively low-communication baseline. tensor can parallelize decode more directly, but currently requires Flash Attention, does not support quantized KV cache, disables auto-fit, and is not implemented for several architectures including Nemotron-H / Nemotron-H-MoE.
The general vLLM parallelism guide likewise notes that pipeline parallelism can be preferable when the GPU-to-GPU path is weak or the split is uneven.
For 3090 specifically, I would still verify:
nvidia-smi topo -m
and actual P2P/NCCL behavior if tensor parallelism matters. NVLink capability is useful, but bridge presence, PCIe topology, BIOS/IOMMU/ACS, runtime build, and backend behavior can still change the effective path.
This is why I like “48GB of placement headroom” better than “one 48GB GPU.”
That headroom can still be extremely useful: less CPU offload, less aggressive quantization, more KV cache/context, higher-precision modality encoders, room for Talker/codec stages, and enough slack to test several layouts instead of immediately fighting OOM.
Qwen3-Omni’s official memory numbers
For the documented Transformers + BF16 + FlashAttention2 path:
| Model |
15s video |
30s |
60s |
120s |
| Qwen3-Omni Thinking |
68.74GB |
77.79GB |
95.76GB |
131.65GB |
| Qwen3-Omni Instruct |
78.85GB |
88.52GB |
107.74GB |
144.81GB |
These numbers show why aggregate 48GB does not imply a BF16 fit. They do not show that quantized Qwen3-Omni cannot run on two consumer GPUs.
The 2×4080 SUPER INT4 example is useful precisely because it got past “does it load?” and exposed a later failure: short speech worked while long speech degraded.
For a companion I would therefore qualify in layers:
load
-> text
-> image
-> audio
-> joint A/V
-> short speech
-> 30–60s speech
-> long speech
-> streaming/interruption
-> long-session memory/latency stability
Model branches I would keep separate instead of making one "best omni model" list
Native omni understanding, text output
The central ~30B branches are:
Both take text/image/audio/video and return text. This is the clean branch if the companion needs to see/hear but speech synthesis can be external.
Nemotron’s ~3B-active figure describes token-time routing/compute; it should not be read as “only 3B parameters need to reside in memory.”
Native speech-output omni
If speech should come from the same end-to-end family, Qwen3-Omni Instruct is the obvious Qwen branch because it includes Talker/Code2Wav.
MGM-Omni-32B is another speech-output branch aimed at longer-horizon speech interaction/generation.
EMOVA-72B is a larger speech/companion-oriented research branch.
Adding native speech changes the runtime/memory/stability problem enough that I would keep this separate from text-output-only omni models.
Keep the preferred language brain and add specialist perception
This became more interesting the more I looked.
LFG-3 is explicitly:
Gemma 4 31B
+ Parakeet audio encoder
+ learned projector
for speech-aware reasoning/conversation.
Ultravox shows the broader modular pattern: a pretrained language backbone plus a Whisper speech encoder and multimodal adapter.
So if the real requirement is:
“I like the way this model thinks/talks, but I want ears,”
the design space includes:
preferred LLM
+ qualified audio frontend
+ external TTS if desired
rather than forcing a complete checkpoint replacement.
The cost is integration/alignment: you cannot attach an arbitrary encoder/projector and expect it to work without training/qualification.
Full-duplex / realtime interaction
I would not infer this from size.
A system can have audio input and speech output but still be turn-based. A natural realtime companion cares about interruption, backchannels, first-audio latency, turn timing, continuous listening, speaking while receiving input, and knowing when not to answer.
MiniCPM-o 4.5 is a useful counterexample: it is only 9B, but its model card explicitly targets continuous audio/video input with concurrent text/speech output and full-duplex interaction.
That does not mean 9B beats 30B at philosophy. It means interaction intelligence is a separate axis from reasoning-model scale.
Why I am not treating 70B as the automatic next step
The public A/V results are useful as a sanity check.
The Daily-Omni leaderboard places Nemotron 3 Nano Omni 30B-A3B and Qwen3-Omni 30B-A3B among the stronger open-weight entries for synchronized audio+visual reasoning. I would not treat that as a general intelligence ranking; it measures a specific capability axis. But it does show modern open ~30B omni systems are serious A/V models, not obviously a dead-end tier.
The video-SALMONN 2 project is also a useful size counterexample. Its author-reported table places its newer 32B Pro above its older 72B 2+ branch on the listed A/V metrics. Again, that does not prove 32B is globally better. It shows training/architecture can matter more than raw size on a target task.
If I wanted to decide whether 70B is worth it, I would use 5–10 real target conversations with:
same system prompt
same prior history
same persona policy
same user messages
adequate generation budget
Then I would judge:
- nuance;
- willingness to disagree;
- correction after challenge;
- overconfidence;
- long-turn consistency;
- instruction adherence;
- conversational feel.
After that I would run the actual modality requirements:
one image-dependent task
one audio-dependent task
one short joint A/V task
one modality-ablation control
Then the decision becomes:
if ~30B is already good enough:
stop scaling and work on perception/memory/voice/interaction
else:
identify the specific failure
test whether a larger branch fixes that failure
That gives the 70B deployment cost a reason.
For the larger branches themselves, I would keep at least two categories separate:
- EMOVA-72B: speech/companion direction;
- video-SALMONN 2+ 72B: A/V-understanding direction.
What I have not found is a simple open model I would summarize as “Qwen3-Omni, same modalities/output/runtime, but 72B.” These are different systems, not one smooth scaling curve.
Deployment traps and the 2×3090 qualification sequence
A few things seem especially easy to misread:
- Aggregate VRAM != unified VRAM. Two cards increase total placement capacity; they do not make every tensor behave as though one device had the sum.
- Checkpoint size != runtime peak memory. Runtime also needs KV cache, workspaces, modality state, encoders/projectors, speech stages, and concurrency headroom.
- Active MoE parameters != resident parameters. A3B describes routed compute, not a 3B memory footprint.
- Quantization is component-scoped. A “4-bit” omni stack may still have encoders, projectors, Talker, codec, or KV cache at a different precision.
- Load success != modality correctness. Each modality needs a prompt whose answer genuinely depends on that modality, ideally plus an ablation.
- Speech output != audio understanding != full-duplex.
- Short speech != stable long speech.
- Thinking budget and chat template are part of evaluation. The same weights can behave very differently with thinking on/off or a different template.
- Personality is a system property. Checkpoint, memory, prompt, session state, voice, disagreement policy, tools, and initiative policy all contribute.
If the real target is 2×3090, I would freeze the exact recipe before making a fit claim:
GPU topology / bridge / driver / CUDA
exact checkpoint + revision
quantization and higher-precision submodules
runtime build
vision/audio/projector files
Talker / Code2Wav / TTS if needed
target context
concurrency
GPU0/GPU1 placement
Then qualify in the cheapest order:
text smoke
-> per-GPU memory ledger
-> image-dependent test
-> audio-dependent test
-> joint A/V
-> modality ablation
-> representative context
-> short/30s/60s/long speech if needed
-> interruption/full-duplex if needed
-> long-running session
For the parallelism itself, I would start with the architecture’s natural split:
Qwen3-Omni full pipeline:
stage placement first
ordinary GGUF decoder/VLM:
layer split first
then, if supported and useful:
tensor parallel as a latency experiment
That minimizes the number of moving parts while still using both cards seriously.
One intermediate branch worth keeping in mind: Qwen3.8-27B
Qwen3.8-27B is interesting if the requirement shifts toward:
strong text
+ image
+ video
+ controllable thinking
- native audio
So it does not replace Qwen3-Omni’s “ears,” but it is a strong intermediate-size text/vision/video branch.
Its local direct-mode run also illustrates why nominal size is a poor speed predictor. Turning thinking off removed reasoning-token overhead, but the dense 27B model still decoded around ~1.7–1.8 tok/s on the single L4 setup.
There are two separate questions:
How many tokens does the policy generate?
How expensive is each generated token?
A sparse ~30B-A3B model can generate each token faster than a dense 27B model while still taking longer to answer if it produces far more reasoning.
If I had to compress all of this into one practical recommendation, it would be:
Try to make a ~30B-class model fail on something you actually care about before paying the cost of the 70B class.
If the ~30B candidate already gives you the depth of conversation you want, then the harder engineering problems become perception, memory, voice, latency, speech stability, interruption, and multi-GPU placement — none of which are automatically fixed by doubling the parameter count.
And if you already like the language behavior of Gemma, I would keep the modular “preferred brain + specialist ears” route open rather than assuming you must trade away the brain you like to get audio.
For the hardware side, I would currently describe 2×3090 as a plausible and materially more flexible platform for quantized ~30B omni experiments, while keeping the exact fit/performance claim conditional on the model, quantization, runtime, modality stack, and stage/layer/tensor placement.