Machinelearningmastery iconMachinelearningmasterySep 17, 2026 ~7 min source read

How to build a multilingual text classifier with Scikit-LLM and open multilingual embeddings

A practical walkthrough that uses a local Ollama server running BGE-M3 to generate multilingual embeddings, then trains a lightweight scikit-learn classifier on top — no per-language models and no paid APIs required.

Multilingual Text Classification with Scikit-LLM and Multilingual Embeddings

Share this story

Send the public story page.

Useful takeaways from this story.

Multilingual LLM embeddings map different languages into a common vector space, so one downstream classifier can handle multiple languages.

You can run the full pipeline locally and for free by using Ollama with an open multilingual model (BGE-M3) and configuring Scikit-LLM to point at the local server.

Use a manageable, language-balanced sample (the example uses 1,000 English and 1,000 Spanish reviews) and random shuffling before splitting into train/test to avoid sampling bias during evaluation.

# Overview This guide explains how to construct a multilingual text classification pipeline without training separate models per language. The approach uses multilingual embeddings produced by an open LLM (BGE-M3) served locally via Ollama, with Scikit-LLM handling the embedding calls and scikit-learn providing the downstream classifier.

# Why use multilingual embeddings

# Environment and tooling The example stack in the article uses:

  • Ollama to host models locally and avoid paid APIs.
  • BGE-M3 as the multilingual embedding model.
  • Scikit-LLM for seamless integration between embeddings and scikit-learn.
  • scikit-learn (logistic regression in the example) for the downstream classifier.
  • Reviews dataset as a real-world evaluation corpus.

The guide includes commands to install dependencies, start Ollama as a background process, pull the BGE-M3 model, and set Scikit-LLM configuration to point at http://localhost:11434/v1/ with a dummy key accepted by the local server.

# Data selection and sampling

# Pipeline steps (high-level)

  • Start Ollama and pull the BGE-M3 model.
  • Configure Scikit-LLM to use the local Ollama endpoint and provide a dummy key.
  • Load and shuffle the multilingual dataset, selecting a balanced sample per language.
  • Generate embeddings for each review using Scikit-LLM calling the local model.
  • Train a simple scikit-learn classifier (logistic regression in the article) on the embeddings.
  • Evaluate on a held-out test split to measure downstream classifier performance.
  • Using a local Ollama instance avoids paid APIs and keeps the solution runnable in notebook environments.

# When this approach fits This pipeline suits projects that need a single classifier across multiple languages and where keeping inference costs and external API dependency low is important. It is also useful when you prefer to run models locally or within controlled environments.

# Limitations and next steps to consider The walkthrough focuses on a minimal, runnable example. For production use, consider embeddings caching, batching, model versioning, and monitoring for embedding drift. You may also scale the training sample and evaluate other classifiers or calibration steps depending on label distribution.

# Bottom line Generate multilingual embeddings once with a multilingual LLM, train a lightweight scikit-learn classifier on those embeddings, and you can classify texts across languages without separate per-language models. The article provides the concrete commands and code snippets to set up Ollama, pull BGE-M3, configure Scikit-LLM, and run the full pipeline on a sample multilingual review dataset.

More context around this story.

300 классов, 400 сэмплов и никакого промптинга: ЛЛМ как энкодер для классификации
Habr iconHabrSep 25, 2026

300 классов, 400 сэмплов и никакого промптинга: ЛЛМ как энкодер для классификации

Когда речь заходит о ЛЛМ и классификации, первая мысль — попросить модель сгенерировать класс через промптирование модели с описанием правил и добавлением примеров. Для небольшого количества классов это работает. Но для 100+ классов (или даже 20-30) — уже нет: контекст раздувается, модель галлюцинирует, а инференс може

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app