The First Instruct-Following Large Language Models for Hungarian

Zijian Győző Yang, Réka Dodé, Gergő Ferenczi, Péter Hatvani, Enikö Héja, Gábor Madarász, Noémi Ligeti-Nagy, Bence Sárossy, Zsófia Szaniszló, Tamás Váradi, Tamás Verebélyi, Gábor Prószéky · 2024

In recent months, large language models have gained significant attention, with companies striving to develop models capable of solving various natural language processing tasks through extensive data training. The release of ChatGPT by OpenAI demonstrated unprecedented capabilities via a multi-step fine-tuning process. For Hungarian, pre-trained large language models include PULl GPT-3SX, PULl GPTrio and in the recent months SambaLingo. In our research, we pre-trained a new large language model based on Llama-2 and inspired by ChatGPT, focuses on fine-tuning with instruction-based prompts. We created a Hungarian prompt dataset and fine-tuned the PULl large language models into instruction-following models. In our research, we discovered that transfer learning allows the model to gain insights from other languages. We found that further pre-training of the language model could leverage valuable knowledge from the originally pre-trained model. Additionally, we can adapt a LLaMA model to another language, such as Hungarian. Our PULl LlumiX models in three Hungarian benchmark could achieve significant better performance. Our instruction model in both HuSST and HuRTE zero-shot competitions could achieve more than 10 accuracy scores. Our further pre-trained Llama-2 model, the PULl LlumiX 32K and the fine-tuned PULl LlumiX 32K Instruct, became state-of-the-art models capable of solving various lanauage technology problems,

Read the paper · More papers on PaperTik