easytools

AI Speech-to-Text Transcriber Beta

Convert audio files to text instantly using the NVIDIA Nemotron 3.5 ASR (0.6B) model.

🔒 100% Private
🗣️ Powered by NVIDIA Nemotron-3.5 ASR (0.6B)
🌍 Supports 40+ Language Locales
💡 Note: This tool is in Beta. It runs on the small NVIDIA Nemotron 3.5 ASR (0.6B) model, so transcription may take longer and might yield slightly inaccurate results.
Use this for both uploaded files and push-to-talk voice capture. Live mode works best with short clips or long press-and-hold speech.
🎙️

Drop your audio or video file here

or click to browse from device (MP3, WAV, M4A, etc.)

Push-to-Talk Voice

Hold Space or press and hold the button while speaking. Release to process the clip.

Idle
Live transcript Release to convert speech into text

Frequently Asked Questions

What model is running behind this tool?

We host the official NVIDIA Nemotron-3.5-ASR-Streaming-0.6b model. It uses the FastConformer architecture optimized for low-latency transcribing, native capitalization, and formatting.

Is my audio data secure?

Yes. Your uploaded files are stored temporarily in memory or a secure `/tmp` directory during processing, and are wiped immediately after the transcription is completed. We do not store or catalog your files.

Does it support long audio files?

Yes. The backend parses files of various durations. However, for large files, conversion and transcription might take a few moments. We limit uploads to a generous size.