VoxMem – An Open-Source Audio Large Model Memory Benchmark from the University of Melbourne and Others

Executive Summary:
VoxMem is an audio large model memory benchmark introduced jointly by the research team from the University of Melbourne and the University of New South Wales. This benchmark pioneers a two-dimensiona...
1. What is VoxMem
VoxMem is an audio large model memory benchmark introduced jointly by the research team from the University of Melbourne and the University of New South Wales. This benchmark pioneers a two-dimensional evaluation framework of "4 types of acoustic evidence × 4 memory operations," deeply integrating four information dimensions—speech semantics, speaker identity, paralinguistic cues, and environmental sounds—with four memory operations: information extraction, cross-session reasoning, temporal tracking, and rejection answering. It establishes a standardized evaluation system covering multi-session histories from 8K to 64K tokens. Systematic testing on 15 mainstream models shows that under a 32K context length, no model achieves an overall accuracy exceeding 40%. While models can remember "what was said," they struggle with answering "who said it" and "how it was said." This finding indicates that merely extending the audio context is insufficient to support long-term memory capabilities in voice interactions.

Technical positioning and domain: VoxMem belongs to the field of large language model evaluation and benchmarking, specifically focusing on assessing the long-term memory capabilities of large audio language models (Large Audio Language Model, LALM). It is the first benchmark framework to systematically incorporate non-text acoustic information into speech memory evaluations.
Development background: This benchmark was jointly developed by the research team from the University of Melbourne and the University of New South Wales. The team has extensive experience in speech processing, natural language understanding, and model evaluation. The motivation for its development stems from the fact that existing speech memory evaluation benchmarks are almost entirely based on transcribed text, neglecting speaker characteristics, tone, environmental sounds, and other "native audio" information that cannot be recovered through transcription, thus creating a significant evaluation blind spot.
Core value: VoxMem systematically addresses the evaluation gap in "native audio memory" for the first time, incorporating non-text dimensions such as "who said it," "how it was said," and "what was heard" into a unified and quantifiable evaluation framework. It provides a comprehensive perspective for understanding the memory capabilities of audio large models and offers reproducible quantitative references for model selection in persistent voice interaction products.
Technical features: VoxMem features a strict leakage prevention construction process and a dual-quality inspection mechanism. By ensuring that neither party in the conversation mentions acoustic clues, using paired version comparisons, and eliminating questions that can be answered solely through transcription, it ensures that the evaluation questions truly test the model's "auditory comprehension ability" rather than its "text reading ability." Its interference conversation design and length nesting control also significantly enhance the reliability and purity of the evaluation results.
2. Key Features
Multisession Speech Memory Evaluation: Under a speech history setup involving multiple sessions and across different time points, the system can test whether large audio language models can remember and trace back previous conversation rounds. The evaluation covers information retrieval within a single session and the integration of information across multiple independent sessions, simulating real-world scenarios where users express the same topic in multiple turns or where the state is continuously updated during human-machine speech interaction.
Four Categories of Acoustic Evidence Coverage: The evaluation examines semantic content (what the user said), speaker identity (determined by voiceprint features), paralinguistic cues (tone, emotion, pauses), and environmental sounds (background noise, ambient sounds). The latter three are "native audio" information that cannot be recovered through speech transcription, fundamentally distinguishing this evaluation from pure text-based memory assessments, ensuring full coverage of the semantic space for audio memory.
Four Categories of Memory Operations: Includes information extraction, cross-session reasoning, temporal evolution tracking, and abstention with reasoning when evidence is insufficient. These four operations cross-reference with the four categories of acoustic evidence, excluding the combination of semantic × extraction, which degrades to pure text retrieval, forming 15 effective evaluation cells and establishing a complete evaluation space.
Controlled Historical Length Mechanism: For the same question, the acoustic evidence remains unchanged across four historical lengths of 8K, 16K, 32K, and 64K tokens. Only irrelevant filler sessions are added to extend the history. This strict nested design isolates the impact of "historical length" on memory performance, eliminating interference from question difficulty and evidence type mixing, enabling orthogonal analysis of length effects.
Leak-Proof Construction and Dual Quality Checks: During the dialogue generation phase, neither the user side nor the assistant side directly mentions speaker identity, tone, or environmental sounds, preventing answers from leaking at the textual level. Additionally, for each round with cues, a paired version without cues is preserved for quality inspection. Through a two-tier quality check process, questions that can be answered solely based on transcribed text are excluded, ensuring that each question truly requires "listening" rather than "reading."
Standardized Open Evaluation System: All code and data are open-sourced, providing standardized running scripts (run_voxmembench.py), scoring scripts (score_voxmembench.py), and an official Leaderboard. It supports API models (e.g., OpenAI, Anthropic) and locally deployed models, ensuring reproducible evaluation processes and comparable results.
High-Fidelity Speech Synthesis and Cue Injection: Data construction uses TTS to synthesize speech, with each user having a fixed voice color. Tone cues are injected through style control, and environmental sounds are mixed in at a 10dB signal-to-noise ratio. This ensures that acoustic cues exist and are identifiable at the physical acoustic level.
3. How to Use
Clone the repository: First, pull the project code from GitHub and run
git clone https://github.com/swagshaw/voxmemin the terminal. Then navigate into the project directory withcd voxmem. Ensure that Git tools are installed on your local machine.Install dependencies: Run
pip install -r requirements.txtto install all Python dependencies. If you want audio files to be automatically decoded into array format, you can also runpip install "datasets[audio]". It is recommended to use Python 3.9 or higher to ensure compatibility of dependencies.Configure API key: If you plan to evaluate API-based models (e.g., GPT-4o-audio-preview), you must first install the corresponding SDK and configure the key. For example, with OpenAI, run
pip install openaiand then set the environment variable usingexport OPENAI_API_KEY=<your-key>(Linux/macOS) orset OPENAI_API_KEY=<your-key>(Windows).Run smoke test: First, use a lightweight configuration to verify that the workflow is functioning properly. Run
python run_voxmembench.py --config 8k_speaker_information --model abstain --allow-abstention --limit 25 --out predictions_smoke.jsonl, then runpython score_voxmembench.py --judge exact-matchto score. Confirm that the abstention layer scores 1.0 on abstention questions and 0.0 on answerable questions, ensuring the evaluation pipeline is working correctly.Formal evaluation: Run
python run_voxmembench.py --config 32k --model openai:gpt-4o-audio-preview --allow-abstention --out predictions_32k.jsonlto evaluate the target model. For local models, use the--adapterparameter to specify a custom adapter function and the--model-dirparameter to point to the weights directory. Optional configurations include the full32kdataset and the32k_speaker_informationsubset.LLM judge scoring: Execute
python score_voxmembench.py --predictions predictions_32k.jsonl --judge openai:gpt-4o-mini --out metrics_32k.json. The LLM judge reads only the question, standard answer, and model response, then determines semantic equivalence for open-ended answers. Finally, it outputs a detailed metrics report in JSON format.
4. Pros and Cons Analysis
| Pros |
|---|
| First-of-its-kind audio-native memory evaluation dimension: For the first time, it incorporates information that cannot be recovered through transcription, such as speaker identity, paralinguistic cues, and environmental sounds, into a unified evaluation framework. This fills the gap left by existing benchmarks that only examine lexical semantic content, giving voice memory evaluation a true "auditory" dimension. |
| Rigorous two-dimensional classification framework: Composed of 4 types of acoustic evidence × 4 memory operations, forming 15 effective evaluation cells, replacing fragmented subjective topic selection methods. This allows for a comprehensive, orthogonal, and systematic breakdown of the model's voice memory capabilities, ensuring complete and non-redundant evaluation dimensions. |
| Complete decoupling of length effects: For the same question, the acoustic evidence remains strictly nested and unchanged across four historical lengths (8K to 64K). Only irrelevant filler conversations are added to the history, cleanly separating the impact of "longer history" from question difficulty and evidence type, ensuring clear and reliable evaluation attribution. |
| Dual anti-leakage mechanisms ensure question purity: Neither party in the conversation mentions acoustic cues, paired version comparisons, and two-tier quality checks eliminate questions that can be answered purely through transcription. Combined with interfering conversations that are topic-similar but have different values, this blocks shortcuts relying on keyword matching, ensuring that the test results effectively reflect auditory memory capabilities. |
5. Comparative Analysis with Similar Tools
| Comparison Dimension | VoxMem | LongMemEval | LoCoMo |
|---|---|---|---|
| Core Modality | Audio-native (user turns are speech, including acoustic cues) | Pure text dialogue history | Pure text dialogue records |
| Evaluation Target | Large audio language models (LALM) | LLM memory systems and long-context models | Long conversation memory models |
| History Length | 8K/16K/32K/64K audio tokens, with identical evidence across four levels for the same question | Approximately 115K tokens (S version), fixed per level | Up to 50 conversation turns, about 7K-10K tokens |
| Memory Dimensions | 4 types of acoustic evidence + 4 memory operations cross-referenced (15 effective cells) | Extraction, multi-session reasoning, temporal, knowledge updating, refusal to answer | Information extraction, temporal reasoning, multi-hop reasoning, knowledge conflict |
| Acoustic Evidence | Full coverage of semantic, speaker, paralinguistic, and environmental sound categories | None, only lexical semantic information | None, only lexical semantic information |
| Interference Design | Similar-topic interfering conversations + length nesting control + dual-quality checks to prevent leakage | Dynamically constructed chat history, simulating real dialogue noise | Mixed construction of temporal interference and topic drift |
| Key Findings | Over 40% model performance at 32K, acoustic memory is significantly weaker than semantic memory | Models lag significantly in knowledge updating and long-history retrieval | Model fidelity and fact-checking capabilities degrade significantly in long conversations |
| Open Source Status | Code under MIT license + data under CC BY-NC, includes Leaderboard | Public benchmark, with open code and data | Public dataset and evaluation scripts |
From the comparison, it is evident that VoxMem is currently the only benchmark that evaluates speech memory from an audio-native perspective. It demonstrates clear differentiation in the completeness of evaluation dimensions and its anti-leakage design. While LongMemEval and LoCoMo are relatively mature in text memory evaluation, they completely ignore the non-transcribable acoustic properties present in speech information.
In practical selection, if the product is oriented toward voice interaction scenarios—such as smart speakers, in-car voice assistants, or voice customer service systems—VoxMem is the benchmark solution for evaluating the memory capabilities of models. Its speaker and paralinguistic dimensions directly reflect key issues in real voice interactions, such as "who said what" and "how the tone changes." If the product primarily involves pure text-based dialogue (such as text-based customer service chatbots or document Q&A systems), the text evaluation frameworks of LongMemEval or LoCoMo provide more detailed coverage in knowledge updating and multi-hop reasoning. For agent systems requiring long-term memory management, the evaluation focus of MemGPT-like frameworks lies in memory hierarchy and storage efficiency, which complements rather than replaces the evaluation goals of VoxMem.
6. Editor's Summary
The launch of VoxMem marks a significant shift in the evaluation of audio large models from a "text-centric" approach to an "audio-native" one. At its core, VoxMem highlights a long-overlooked fundamental truth in speech memory: the information carried by speech extends far beyond just lexical content. The speaker's identity, tone variations, and environmental sounds collectively form the complete information stream of human auditory memory. By conducting orthogonal evaluations of four types of acoustic evidence and four memory operations within a unified framework, VoxMem not only fills the evaluation gaps in existing benchmarks but also clearly defines the actual capability boundaries of current audio models through experimental results showing that none of the 15 models tested achieved an overall accuracy exceeding 40% at a 32K context length.
From a technical implementation perspective, the reliability of VoxMem's evaluations is built on a rigorous construction methodology. The three-stage data generation process ensures that acoustic cues physically exist and are not leaked through text; the use of an interference session paired with a text classifier that achieves only 51%-56% accuracy effectively rules out speculative paths based on keyword matching; and the length nesting control cleanly isolates the effects of historical context length. These design choices give VoxMem's evaluation results high internal validity, with each data point clearly attributable to a specific memory operation and type of acoustic evidence.
In terms of practical value, VoxMem provides a standardized tool for testing memory capabilities in real-world audio products such as voice assistants, in-car interaction systems, multi-speaker customer service platforms, and companion robots. Its multi-session design closely mirrors actual human-computer interaction patterns, and the four types of acoustic evidence correspond to typical interference sources encountered during product operation. The evaluation results have direct reference value for product selection. The project's code and data are fully open-sourced, with a Leaderboard provided, and the evaluation protocol is standardized, which facilitates the development of a community-driven, continuous evaluation ecosystem. For audio large model development teams, VoxMem's evaluation results can also serve as diagnostic indicators for model iteration, promoting specialized optimization in weaker areas such as speaker binding and emotional state tracking. As multimodal large models rapidly integrate into voice interaction scenarios, the "acoustic memory" evaluation dimension defined by VoxMem is expected to become a foundational component of the assessment framework in this field.
7. Application Scenarios
Voice Assistant/Smart Speaker Long-Term Memory Evaluation: When vendors are selecting or evaluating the underlying models for voice assistants, they can use VoxMem to uniformly test whether candidate models can remember user preferences, voice characteristics of family members, and previous commitments across sessions, avoiding user experience issues such as "forgetting what was said last week." By comparing the performance of multiple models across 15 evaluation cells, developers can identify the strengths and weaknesses of each model in speaker binding and cross-session reasoning, aiding in the decision-making process.
In-Vehicle and Wearable Voice Interaction Evaluation: In-car environments are noisy and users' tones can vary significantly, and wearable devices are also subject to external sound interference. This places high demands on the noise-robust memory capabilities of voice models. VoxMem's environmental sound and paralinguistic cues dimensions can specifically test the reliability of a model's memory in real-world noisy scenarios, helping developers assess the actual performance of models in complex acoustic environments and optimize their noise robustness accordingly.
Multi-Speaker Customer Service and Meeting System Quality Assurance: In scenarios such as call centers and meeting transcription systems, it is essential to correctly bind "which sentence was said by whom." Incorrect speaker attribution can directly lead to misjudgment of responsibility and distorted transcription materials. VoxMem's speaker evidence dimension, combined with information extraction tasks, can be directly used to diagnose "mixing up identities" issues in such systems, enabling systematic analysis of speaker binding attribution errors to identify weak links in the speaker recognition pipeline.
Emotional Memory Validation for Companion Robots/Aging Care Products: Information such as matters mentioned by elderly users last week with sighs, or the fatigue sensed in their tone, can only be carried by paralinguistic cues. Text transcription would completely lose these details. VoxMem's temporal tracking dimension can assess whether companion products can detect changes in the user's emotional state across time, providing a quantifiable evaluation standard for the "emotional continuity" of aging care products, and guiding product iteration and development in the direction of emotional memory.
8. FAQ
Q: What is the fundamental difference between VoxMem and existing audio understanding benchmarks?
A: Existing audio understanding benchmarks primarily evaluate a model's immediate comprehension ability — transcribing, answering questions, or summarizing content after listening to an audio clip, with the focus on "how much was understood at the moment." VoxMem evaluates the model's long-term memory capability across sessions — whether the model can accurately recall specific information from previous audio clips even after experiencing multiple sessions and a large amount of irrelevant information. This includes non-textual aspects such as who spoke, the tone, and background sounds. In short, the former addresses the question of "whether the model can understand," while the latter addresses "whether the model can remember," with the memory content extending to native audio information that cannot be recovered through transcription.
Q: Why is the "semantics × information extraction" combination excluded from the evaluation framework?
A: Combining semantic evidence (what the user said) with information extraction (directly extracting answers from a single session) is essentially a pure text retrieval task — the model only needs to transcribe the audio into text and then locate the answer within the text, without involving any auditory memory capability. After excluding this degenerate combination, the remaining 15 cells all require the model to rely on non-textual information or cross-session integration to some extent, ensuring that each evaluation dimension truly tests "acoustic memory" rather than text retrieval ability.
Q: How is the possibility of models cheating by transcribing audio to text and then finding answers prevented?
A: VoxMem's dataset construction employs a dual safeguard design. First, during the dialogue generation phase, it explicitly requires that neither the user's nor the assistant's dialogue text mention speaker identity, tone, or environmental sounds, making it impossible to obtain the correct answer by simply reading the transcribed text. Second, the interfering session and the evidence session have similar topics but different specific values, and a classifier based solely on text (with an accuracy of only 51-56%) is used to verify the effectiveness of the interference, ensuring that the model cannot find answers through superficial topic matching and must truly "hear" the clues at the audio level.
Q: Does the historical length setting in the evaluation significantly affect the model's API cost and evaluation time?
A: Yes, it has a significant impact. Among the four length settings, the 64K setting includes a large amount of padding audio in each evaluation sample, resulting in significantly increased transmission and inference time. The token consumption for API calls is calculated based on the audio length, and the cost of running the full evaluation set will be several times higher than that of the 8K setting. It is recommended to first use a single subset (e.g., speaker_information) of the 8K or 16K setting to verify the process, then proceed with the full 32K evaluation. If long-context capabilities need to be assessed, the 64K setting can be run selectively to control the overall evaluation cost.
Q: How can an open-source audio model deployed locally be integrated into the VoxMem evaluation process?
A: A locally deployed model must use the --adapter parameter to specify a custom adapter function, encapsulating the model's inference interface into the unified calling format required by VoxMem. At the same time, the --model-dir parameter should point to the model's weight directory. The official repository provides example adapter code covering two mainstream deployment methods: loading from HuggingFace Transformers and inference with vLLM. If the model supports an OpenAI-compatible audio input format, it can also be directly routed to a local gateway service using --model openai:<model-name>.
Q: How is the reliability of the LLM judge scoring ensured?
A: During the scoring phase, the judge model specified by the --judge parameter determines the semantic equivalence between open-ended answers and standard answers, rather than relying on simple string matching. The VoxMem paper conducted calibration experiments to ensure the correlation between judge scores and human scores, and the selected judge models (e.g., gpt-4o-mini) demonstrate stable performance in semantic equivalence judgment. Additionally, the evaluation script supports the exact-match mode as a control, allowing researchers to choose between exact matching and semantic equivalence based on their specific requirements.
9. Project Links
- Project Website: https://swagshaw.github.io/voxmem/
- GitHub Repository: https://github.com/swagshaw/voxmem
- HuggingFace Dataset and Model Library: https://huggingface.co/datasets/AudioMemory/voxmembench
- arXiv Technical Paper: https://arxiv.org/pdf/2609.32607
Related AI Model Articles
Vidu Q4 Preview: A Masterful Demonstration of Expressive Performance and Multi-Subject Consistency in ShengShu Technology's Flagship Video Generation Model
Vidu Q4 Preview is the first publicly available preview version of ShengShu Technology's next-generation flagship video generation model, launched on October 7, 2026. It emphasizes "expressive perform...

EmbeddingGemma 2 – Review of Google's Open-Source Native Multimodal Embedding Model
EmbeddingGemma 2 is an open-source native multimodal Embedding model developed by Google, built upon the Gemma 4 architecture. It unifies five modalities—text, code, images, video, and audio—into a si...
MAI-Transcribe-2-Streaming: In-Depth Evaluation of Microsoft's Real-Time Streaming Speech-to-Text Model
MAI-Transcribe-2-Streaming is Microsoft's first real-time streaming speech-to-text model, capable of continuously outputting text while speech is being delivered. It supports 60 languages and automati...

Index-Translate – Bilibili's Open-Source Multilingual Translation Model Family
Index-Translate is an open-source multilingual translation model family developed by Bilibili's Index LLM team. It is built upon the Qwen3.5 foundation and opens its weights under the Apache-2.0 licen...
© All Rights Reserved. Some content on this site is partially generated by AI with human review.
