DeepSeek-OCR
DeepSeek logo

DeepSeek-OCR

DeepSeek
DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical character recognition (OCR) and “contextual optical compression.” The model is designed to explore the limits of compressing contextual information from images, efficiently processing documents and converting them into structured text formats such as Markdown. The model requires an image as input.

Pricing

  • Input Tokens: $0.02 /M tokens
  • Output Tokens: $0.02 /M tokens

Input Modalities

  • Text
  • Vision

Output Modalities

  • Text

Providers

Baidu baidu-deepseek-ocr
Pricing$0.0411$0.1644
Context32K
Max output28K
Latency1.2S
Throughput89.3TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Siliconflow deepseek-ai/DeepSeek-OCR
Pricing$0.02$0.02
Context8K
Max output0
Latency-
Throughput-
Uptime
100.00% uptime 2 days ago
0.00% uptime yesterday
100.00% uptime today

Performance for DeepSeek-OCR

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="DeepSeek-OCR",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is DeepSeek-OCR?

DeepSeek-OCR is a vision-language model launched by DeepSeek AI, focusing on optical character recognition (OCR) and “contextual optical compression.” The model is designed to explore the limits of compressing contextual information from images, efficiently processing documents and converting them into structured text formats such as Markdown. The model requires an image as input.