Viking-33B Open Source Model - Supports Multilingual Processing and Has Code Comprehension and Generation Capabilities

Viking 33B

Developed by LumiOpen

Viking 33B is a 33 billion parameter decoder-only Transformer model supporting Finnish, English, and multiple Nordic languages, with capabilities in code understanding and generation.

Large Language Model

Transformers

Supports Multiple LanguagesOpen Source License:Apache-2.0 #Nordic multilingual processing #33 billion parameter large model #Code generation and comprehension

Downloads 1,030

Release Time : 2/20/2024

Model Overview

A multilingual large model focused on Nordic languages and code processing, jointly developed by institutions including the University of Turku, Finland, under the Apache 2.0 open-source license.

Model Features

Nordic language optimization

Specially optimized for low-resource Nordic languages like Finnish, with basic translation capabilities

Integrated coding capabilities

Training data includes programming code, supporting code understanding and generation tasks

Supercomputer-level training

Trained on the LUMI supercomputer using 1024 AMD GPUs

Progressive release

Training checkpoints released at every 100 billion token interval

Model Capabilities

Multilingual text generation

Cross-language translation

Code generation

Code completion

Technical documentation processing

Use Cases

Language services

Nordic language translation

Basic translation between similar languages like Swedish, Norwegian, and Danish

Finnish text generation

Generating grammatically correct Finnish text

Programming assistance

Code autocompletion

Completing code snippets based on context

🚀 Viking 33B

Viking 33B is a 33B parameter decoder-only transformer. It is pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code, aiming to provide language processing capabilities in multiple Nordic languages and programming contexts.

🚀 Quick Start

Viking 33B is a powerful model with a wide range of applications. For more detailed usage, please refer to the following sections.

✨ Features

Multilingual Capability: Fluent in Finnish, English, Scandinavian languages, and capable of basic translation between them.
Code Understanding and Generation: Can understand and generate code.
Open Source: A fully open source model available under the Apache 2.0 License.

📦 Installation

No specific installation steps are provided in the original document.

💻 Usage Examples

Basic Usage

branch = "200B"
model = transformers.AutoModelForCausalLM.from_pretrained(
    "LumiOpen/Viking-33B",
    torch_dtype=torch.bfloat16,
    revision=branch,
)

📚 Documentation

Model Family

Viking is the second set of models released by LumiOpen and is available at 3 parameter counts:

Model Overview

NOTE: Viking is a base model which needs further fine tuning for most use cases.

Viking is a generative pretrained transformer using a LLaMA-like GPT architecture, and makes use of rotary positional embeddings and flash attention.

Property	Details
n_parameters	33B
n_layers	56
n_heads	56
d_model	7168
vocab_size	131072
sequence_length	4096

Training

Viking 33B was trained on the LUMI supercomputer, using 1024 AMD MI250X GPUs. Each MI250X GPU has two Graphics Complex Dies (GCDs) for a world size of 2048 during training, using activation checkpointing, a micro batch size of 1, gradient accumulation of 16, and a 3D parallelism strategy of TP=4, PP=4, DP=128.

Training began in September 2023 using a custom fork of the Megatron-Deepspeed framework.

Training Hyperparameters

Property	Details	Comment
Precision	bfloat16
Optimizer	AdamW
Learning rate	3e-4	10B tokens warm-up, cosine decay to 3e-5
Weight decay	1e-1
Batch size	1024	1024 samples x 4096 tokens = 4194304 tokens

Tokenizer

Viking uses a custom 128K Bloom tokenizer trained on the same English, Finnish, Swedish, Danish, Norwegian, Icelandic and code dataset used to train the model.

Dataset

Viking is being trained on a 2 trillion token mixed dataset of English, Finnish, Swedish, Danish, Norwegian, Icelandic and code. Full details will be published soon.

Evaluation Results

Full evaluation results will be published with the final model.

Training Checkpoints

Training checkpoints are available as branches in the repository. Checkpoints will be released roughly every 100B tokens. The main branch will always point to the latest checkpoint. The following checkpoints are available:

Ethical Considerations and Limitations

Viking 33B is a release of a partially trained model, and special care should be taken when using any output.

Viking is an advanced language model, primarily optimized for English, Finnish, Swedish, Norwegian, Danish, Icelandic and code, with no meaningful proficiency in any other languages. As with most AI-driven systems, Viking is a product of the vast data it has been trained on, which may reflect the imperfections, biases, and idiosyncrasies of the wider web. Viking may, at times, produce outputs that can be considered inaccurate, prejudiced, or controversial. Users and developers engaging with Viking should exercise discretion and consider additional evaluation and customization to ensure the model's responses align with their specific needs and ethical standards.

🔧 Technical Details

The model uses a LLaMA-like GPT architecture, rotary positional embeddings and flash attention. It was trained on the LUMI supercomputer with specific parallelism strategies and hyperparameters.

📄 License

Viking is released under the Apache 2.0 license.

📚 Citation Information

@misc {lumiopen_2025,
	author       = { Luukkonen, Risto and Burdge, Jonathan and Zosa, Elaine and Komulainen, Ville and Sarlin, Peter and Pyysalo, Sampo },
	title        = { Viking: A Family of Nordic LLMs },
	year         = 2025,
	url          = { https://huggingface.co/LumiOpen/Viking-33B },
	doi          = { 10.57967/hf/4885 },
	publisher    = { Hugging Face }
}

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご