Llama 3.2 90B Instruct

Large multimodal language model from Meta for image analysis and text generation.

Overview

Llama 3.2 90B Instruct

Description of the Llama 3.2 90B Instruct Neural Network

Llama 3.2 90B Instruct is a large multimodal language model developed by Meta. It can process both text and images simultaneously, opening up broad opportunities for visual recognition tasks, analyzing graphic content, and automatically generating image captions.

The model supports a context of up to 128,000 tokens, allowing it to work with large amounts of data in a single request. This is especially important for tasks that require analyzing long documents or large images. Llama 3.2 90B Instruct is optimized for deployment not only on server infrastructure but also on edge and mobile devices, making it accessible for a wide range of use cases.

The model was released on September 25, 2024, and is distributed under the llama3_2 license. It is available both via API and for local deployment, including through the Hugging Face repository.

Multimodal Capabilities

Llama 3.2 90B Instruct combines text and image processing. The model can “see” images, analyze their content, answer questions about images, and generate text descriptions. This makes it a versatile tool for visual AI.

Performance and Benchmarks

In tests, the model shows strong results: AI2D — 92.3%, DocVQA — 90.1%, VQAv2 — 78.1%. The average score across the benchmark suite is 71.3%, confirming its competitiveness among large language models.

Llama 3.2 90B Instruct Specifications

CharacteristicValue
TypeMultimodal language model
Parameters90.0 billion
Context128,000 tokens
Release dateSeptember 25, 2024
Average score71.3%
Licensellama3_2
DeveloperMeta
Price per input (1M tokens)$1.20
Price per output (1M tokens)$1.20
Max input tokens128,000
Max output tokens128,000
Supported capabilitiesFunction Calling, Structured Output, Code Execution, Web Search, Batch Inference, Fine-tuning

Who is the Llama 3.2 90B Instruct neural network suitable for?

Developers and Machine Learning Engineers

The model suits specialists who build computer vision applications, chatbots, document analysis systems, or content generation tools. Thanks to support for Function Calling, Structured Output, and Code Execution, Llama 3.2 90B Instruct can be integrated into complex software products.

Researchers and ML Enthusiasts

The open license and availability of weights on Hugging Face make the model interesting for research projects. The ability to fine-tune and perform batch inference allows the model to be adapted to specific tasks.

Teams Working with Visual Content

Llama 3.2 90B Instruct will be useful for those automating image processing: describing photos, extracting text from documents, analyzing charts and diagrams, and recognizing objects.

How to use the Llama 3.2 90B Instruct neural network?

Via API

The model is available through an API. The model page provides links to API documentation, allowing quick integration. Usage costs $1.20 per 1 million tokens for both input and output.

Local Deployment

For those who want to work with the model on their own servers or devices, a Hugging Face repository is available where the model weights can be downloaded. Llama 3.2 90B Instruct is optimized to run on edge and mobile devices, allowing deployment even in environments with limited computing resources.

Integration into Applications

The model supports many advanced capabilities: Function Calling, Structured Output, Code Execution, Web Search, and Batch Inference. This makes it a flexible tool for embedding into existing systems.

Key Features of Llama 3.2 90B Instruct

Multimodality

The model supports visual recognition, reasoning about image content, and automatic caption generation. It can process both text queries and graphical data within a single session.

Long Context

A context length of 128,000 tokens allows processing large volumes of information — long documents, entire books, multi-page PDFs, or series of images — without splitting the request into parts.

Functional Capabilities

Llama 3.2 90B Instruct supports function calling, allowing the model to interact with external tools and APIs. Structured Output ensures data is returned in a specified format. The ability to execute code and search the web expands possible use cases.

Fine-tuning and Batch Processing

The model supports fine-tuning, allowing it to be adapted to specific tasks and domains. Batch Inference enables processing many requests simultaneously, reducing latency and costs.

Advantages of Llama 3.2 90B Instruct

High Benchmark Performance

The model demonstrates strong results in image and document understanding tasks, such as AI2D (92.3%), DocVQA (90.1%), and VQAv2 (78.1%). This confirms its ability to accurately analyze visual information.

Efficient Edge Operation

Unlike many large models, Llama 3.2 90B Instruct is optimized for deployment on edge and mobile devices. This lowers infrastructure requirements and allows the model to be used in offline scenarios.

Affordable Cost

A price of $1.20 per 1 million tokens for both input and output is competitive for a model with 90 billion parameters. This pricing makes it accessible for large-scale use.

Deployment Flexibility

The model can be used via API, deployed locally, or accessed through Hugging Face. Support for many advanced features (Function Calling, Structured Output, Code Execution, Web Search, Batch Inference, Fine-tuning) allows it to be adapted to a wide variety of tasks.

Disadvantages of Llama 3.2 90B Instruct

High Resource Requirements for Local Deployment

Despite optimization for edge devices, a model with 90 billion parameters requires significant computing power and memory for full local operation. Serious hardware resources are needed to run it properly on standard devices.

Limited Information on Regional Availability

The source data does not specify regional restrictions on model usage. This may create difficulties for users in certain countries or regions where access to Meta models may be limited.

What tasks does Llama 3.2 90B Instruct solve?

Visual Recognition and Image Analysis

The model can recognize objects, scenes, texts, and other elements in images. This allows it to be used for automating visual content processing across various fields.

Reasoning About Image Content

Llama 3.2 90B Instruct not only recognizes objects but can also draw logical conclusions based on what it sees: answering questions, comparing elements, and interpreting scenes.

Image Captioning

One of the model’s key tasks is automatically generating text descriptions for images. This can be applied in assistive systems for the visually impaired, content management, and cataloging.

Generative Tasks

The model can also handle purely text-based tasks: writing texts, answering questions, summarization, and code generation. This makes it a versatile tool not limited to image processing.

Llama 3.2 90B Instruct Pricing

The cost of using the model is $1.20 per 1 million tokens for both input (incoming data) and output (generated text). Thus, the pricing is symmetric — the same rate applies to user requests and model responses.

This pricing approach is standard for many LLMs and makes it easy to calculate costs depending on usage volume. For large projects with high load, the model supports batch inference, which can reduce the final cost for mass requests.

Llama 3.2 90B Instruct Terms of Use

The model is distributed under the llama3_2 license developed by Meta. This license imposes certain conditions on commercial and non-commercial use, including restrictions related to deployment scale and potential areas of application.

Developers are advised to carefully review the license text before using the model, especially if commercial deployment or fine-tuning on proprietary data is planned. The source data does not specify detailed restrictions, so current terms should be checked on the official Meta and Hugging Face pages.

Llama 3.2 90B Instruct Availability

The model is available for download and use through the Hugging Face platform, where its weights are published. It can also be deployed on edge and mobile devices thanks to optimization carried out by the developers.

The model page does not specify regional restrictions. Users in different countries should independently check the availability of Meta services in their region. For API use, it is also necessary to review the provider’s terms of service.

How Llama 3.2 90B Instruct Differs from Alternatives

Multimodality with a Large Context

Among alternatives such as Llama 3.2 11B Instruct, Gemma 3 27B, Mistral Small 3.1 24B Instruct, or GPT OSS 20B, Llama 3.2 90B Instruct stands out with the largest size (90 billion parameters) and support for a 128,000-token context. This allows the model to process significantly more data in a single request than many competitors.

Balanced Cost

At $1.20 per 1 million tokens, the model sits in the mid-price range for its size. Some smaller alternatives may be cheaper but less performant, while larger competitor models may cost more.

Optimization for Edge Devices

Not all large multimodal models are optimized to run on edge and mobile devices. Llama 3.2 90B Instruct was designed with this requirement in mind, setting it apart from many alternatives focused exclusively on server deployment.

Support for a Broad Set of Features

Unlike some alternatives, Llama 3.2 90B Instruct supports all key modern capabilities: Function Calling, Structured Output, Code Execution, Web Search, Batch Inference, and Fine-tuning. This makes it a universal solution that does not require additional tools for integration.

Conclusion

Llama 3.2 90B Instruct is a powerful multimodal model from Meta with a 128,000-token context, optimized for edge and mobile devices. It achieves strong results in benchmarks such as AI2D (92.3%), DocVQA (90.1%), and VQAv2 (78.1%), and offers a broad range of supported capabilities, including fine-tuning and batch processing. Usage costs $1.20 per 1 million tokens for both input and output. The model is available via Hugging Face and API, making it a flexible tool for visual recognition, image analysis, and text generation tasks.

Image analysis
Image caption generation
visual recognition
Processing large volumes of text

Frequently asked questions

Llama 3.2 90B Instruct — Review of Meta's Multimodal Neural Network