Model Details Meta developed and released the Meta Llama 3 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8 and 70B sizes.

Meta-Llama-3-70B-Instruct

Meta-Llama-3-8B-Instruct

This is the instruction fine-tuned version of Mixtral-8x22B - the latest and largest mixture of experts large language model (LLM) from Mistral AI. This state of the art machine learning model uses a mixture 8 of experts (MoE) 22b models. During inference 2 experts are selected. This architecture allows large models to be fast and cheap at inference.

Mixtral-8x22B-Instruct-v0.1

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to those leading proprietary models.

WizardLM-2-8x22B

WizardLM-2 7B is the smaller variant of Microsoft AI's latest Wizard model. It is the fastest and achieves comparable performance with existing 10x larger open-source leading models

WizardLM-2-7B

Zephyr 141B-A35B is an instruction-tuned (assistant) version of Mixtral-8x22B. It was fine-tuned on a mix of publicly available, synthetic datasets. It achieves strong performance on chat benchmarks.

zephyr-orpo-141b-A35b-v0.1

Gemma is an open-source model designed by Google. This is Gemma 1.1 7B (IT), an update over the original instruction-tuned Gemma release. Gemma 1.1 was trained using a novel RLHF method, leading to substantial gains on quality, coding capabilities, factuality, instruction following and multi-turn conversation quality.

gemma-1.1-7b-it

DBRX is an open source LLM created by Databricks. It uses mixture-of-experts (MoE) architecture with 132B total parameters of which 36B parameters are active on any input. It outperforms existing open source LLMs like Llama 2 70B and Mixtral-8x7B on standard industry benchmarks for language understanding, programming, math, and logic.

dbrx-instruct

Mixtral is mixture of expert large language model (LLM) from Mistral AI. This is state of the art machine learning model using a mixture 8 of experts (MoE) 7b models. During inference 2 expers are selected. This architecture allows large models to be fast and cheap at inference. The Mixtral-8x7B outperforms Llama 2 70B on most benchmarks.

Mixtral-8x7B-Instruct-v0.1

The Mistral-7B-Instruct-v0.2 Large Language Model (LLM) is a instruct fine-tuned version of the Mistral-7B-v0.2 generative text model using a variety of publicly available conversation datasets.

Mistral-7B-Instruct-v0.2

LLaMa 2 is a collections of LLMs trained by Meta. This is the 70B chat optimized version. This endpoint has per token pricing.

Llama-2-70b-chat-hf

The Dolphin 2.6 Mixtral 8x7b model is a finetuned version of the Mixtral-8x7b model, trained on a variety of data including coding data, for 3 days on 4 A100 GPUs. It is uncensored and requires trust_remote_code. The model is very obedient and good at coding, but not DPO tuned. The dataset has been filtered for alignment and bias. The model is compliant with user requests and can be used for various purposes such as generating code or engaging in general chat.

dolphin-2.6-mixtral-8x7b

A Mythomax/MLewd_13B-style merge of selected 70B models  A multi-model merge of several  LLaMA2 70B finetunes for roleplaying and creative work. The goal was to create a model that combines creativity with intelligence for an enhanced experience.

lzlv_70b_fp16_hf

OpenChat is a library of open-source language models that have been fine-tuned with C-RLFT, a strategy inspired by offline reinforcement learning. These models can learn from mixed-quality data without preference labels and have achieved exceptional performance comparable to ChatGPT. The developers of OpenChat are dedicated to creating a high-performance, commercially viable, open-source large language model and are continuously making progress towards this goal.

openchat_3.5

LLaVa is a multimodal model that supports vision and language models combined.

llava-1.5-7b-hf

Latest version of the Airoboros model fine-tunned version of llama-2-70b using the Airoboros dataset. This model is currently running jondurbin/airoboros-l2-70b-2.2.1 

airoboros-70b

SDXL consists of an ensemble of experts pipeline for latent diffusion: In a first step, the base model is used to generate (noisy) latents, which are then further processed with a refinement model (available here: https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0/) specialized for the final denoising steps. Note that the base model can be used as a standalone module.

sdxl

Llama 2 is a collection of pretrained and fine-tuned generative text models ranging in scale from 7 billion to 70 billion parameters. This is the repository for the 7B fine-tuned model, optimized for dialogue use cases and converted for the Hugging Face Transformers format. 

Llama-2-7b-chat-hf

Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification.

whisper-large

BGE embedding is a general Embedding Model. It is pre-trained using retromae and trained on large-scale pair data using contrastive learning. Note that the goal of pre-training is to reconstruct the text, and the pre-trained model cannot be used for similarity calculation directly, it needs to be fine-tuned

bge-large-en-v1.5

You can use cURL or any other http client to run inferences:

```bash
curl -X POST \
    -H "Authorization: bearer $(deepctl auth token)"  \
    -F 'instance_prompt=a photo of nkb person'  \
    -F 'class_prompt=a photo of a person'  \
    -F instance_data=@my_archive.zip  \
    -F output_model_name=my_username/my_model_name  \
    'https://api.deepinfra.com/v1/inference/deepinfra/dreambooth'
```

which will give you back something similar to:

```json
{
  "model_name": "my_user_name/my_cool_model",
  "version": "90lkjfslkf5lkdjflskjfsdkl",
  "final_loss": 0.123,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0
  }
}

```


You can use our command-line tool [deepctl](/docs/getting-started) to run
inferences:

```bash
deepctl infer \
    -m 'deepinfra/dreambooth'  \
    -i 'instance_prompt=a photo of nkb person'  \
    -i 'class_prompt=a photo of a person'  \
    -i instance_data=@my_archive.zip  \
    -i output_model_name=my_username/my_model_name
```

which will give you back something similar to:

```json
{
  "model_name": "my_user_name/my_cool_model",
  "version": "90lkjfslkf5lkdjflskjfsdkl",
  "final_loss": 0.123,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0
  }
}

```


The prompt used to describe your training images. Should be something like:'a photo of [identifier] [class noun]'. where [identifier] should be something unique

instance_prompt

The prompt used to describe your class. Should be something like:'a photo of a [class noun]'. where [class_noun] describe a class of instance

class_prompt

zip file containing images of instance you want to train on

instance_data

The name of the model to be saved. Must start with your github username

output_model_name

The maximum number of training steps to run

max_train_steps

train_batch_size

Whether to center crop the images before resizing to resolution

center_crop

Initial learning rate (after the potential warmup period) to use.

learning_rate

The scheduler type to use during training.

lr_scheduler

Number of steps for the warmup in the lr scheduler

lr_warmup_steps

Whether or not to use 8-bit Adam from bitsandbytes.

use_8bit_adam

The beta1 parameter for the Adam optimizer.

adam_beta1

The beta2 parameter for the Adam optimizer.

adam_beta2

adam_weight_decay

adam_epsilon

max_grad_norm

Whether or not to use gradient checkpointing to save memory at the expense of slower backward pass.

gradient_checkpointing

The webhook to call when inference is done, by default you will get the output in the response of your inference request

deepinfra/dreambooth

Input

Output