Skip to main content

Tokenization

OpenVINO GenAI provides a way to tokenize and detokenize text using the ov::genai::Tokenizer class. The Tokenizer is a high level abstraction over the OpenVINO Tokenizers library.

It can be initialized from the path, in-memory IR representation or obtained from the ov::genai::LLMPipeline object.

import openvino_genai as ov_genai

# Initialize from the path
tokenizer = ov_genai.Tokenizer(models_path)

# Or get tokenizer instance from LLMPipeline
pipe = ov_genai.LLMPipeline(models_path, "CPU")
tokenizer = pipe.get_tokenizer()

Tokenizer has encode() and decode() methods that support the following arguments: add_special_tokens, skip_special_tokens, pad_to_max_length, max_length. The Tokenizer can also be tuned for concurrent execution. Below are examples of how to configure these options.

Configure tokenizer streams for concurrent execution

Tokenizer constructors accept OpenVINO compile properties. For workloads that run tokenization from multiple threads, you can increase the number of tokenizer streams to improve throughput.

Pass OpenVINO properties when constructing Tokenizer:

import openvino_genai as ov_genai
# Configure multiple streams for concurrent encode/decode calls.
tokenizer = ov_genai.Tokenizer(models_path, **{"NUM_STREAMS": 4,"INFERENCE_NUM_THREADS": 16})

Choose the stream count based on CPU resources and expected parallel tokenization load.

Example - disable adding special tokens:

tokens = tokenizer.encode("The Sun is yellow because", add_special_tokens=False)

The encode() method returns a TokenizedInputs object containing input_ids and attention_mask, both stored as ov::Tensor. Since ov::Tensor requires fixed-length sequences, padding is applied to match the longest sequence in a batch, ensuring a uniform shape. Also resulting sequence is truncated by max_length. If this value is not defined by used, it's is taken from the IR.

Both padding and max_length can be controlled by the user. If pad_to_max_length is set to true, then instead of padding to the longest sequence it will be padded to the max_length.

Example - control padding:

import openvino_genai as ov_genai

tokenizer = ov_genai.Tokenizer(models_path)
prompts = ["The Sun is yellow because", "The"]

# Since prompt is definitely shorter than maximal length (which is taken from IR) will not affect shape.
# Resulting shape is defined by length of the longest tokens sequence.
# Equivalent of HuggingFace hf_tokenizer.encode(prompt, padding="longest", truncation=True)
tokens = tokenizer.encode(["The Sun is yellow because", "The"])
# or is equivalent to
tokens = tokenizer.encode(["The Sun is yellow because", "The"], pad_to_max_length=False)
print(tokens.input_ids.shape)
# out_shape: [2, 6]

# Resulting tokens tensor will be padded to 1024, sequences which exceed this length will be truncated.
# Equivalent of HuggingFace hf_tokenizer.encode(prompt, padding="max_length", truncation=True, max_length=1024)
tokens = tokenizer.encode(["The Sun is yellow because",
"The"
"The longest string ever" * 2000], pad_to_max_length=True, max_length=1024)
print(tokens.input_ids.shape)
# out_shape: [3, 1024]

# For single string prompts truncation and padding are also applied.
tokens = tokenizer.encode("The Sun is yellow because", pad_to_max_length=True, max_length=128)
print(tokens.input_ids.shape)
# out_shape: [1, 128]