Transformers documentation
GPT-NeoX-Japanese
This model was contributed to Hugging Face Transformers on 2022-09-14.
GPT-NeoX-Japanese, a Japanese language model based on GPT-NeoX. Japanese uses three types of characters (hiragana, katakana, kanji) and has a huge vocabulary. This model uses BPEEncoder V2, a sub-word tokenizer to handle the different characters.
The model also removes some bias parameters for better performance.
You can find all the original GPT-NeoX-Japanese checkpoints under the ABEJA organization.
This model was contributed by Shinya Otani, Takayoshi Makabe, Anuj Arora, and Kyo Hattori from ABEJA, Inc..
Click on the GPT-NeoX-Japanese models in the right sidebar for more examples of how to apply GPT-NeoX-Japanese to different language tasks.
The example below demonstrates how to generate text with Pipeline or the AutoModel, and from the command line.
from transformers import pipeline
pipeline = pipeline(task="text-generation",
model="abeja/gpt-neox-japanese-2.7b", device=0)
pipeline("人とAIが協調するためには、")Quantization reduces the memory burden of large models by representing the weights in a lower precision. Refer to the Quantization overview for more available quantization backends.
The example below uses bitsandbytes to only quantize the weights to 4-bits.
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype="float16"
)
model = AutoModelForCausalLM.from_pretrained(
"abeja/gpt-neox-japanese-2.7b",
quantization_config=quantization_config,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("abeja/gpt-neox-japanese-2.7b")
input_ids = tokenizer.encode("人とAIが協調するためには、", return_tensors="pt").to(model.device)
output = model.generate(input_ids)
print(tokenizer.decode(output[0], skip_special_tokens=True))Use the AttentionMaskVisualizer to better understand what tokens the model can and cannot attend to.
from transformers.utils.attention_visualizer import AttentionMaskVisualizer
visualizer = AttentionMaskVisualizer("abeja/gpt-neox-japanese-2.7b")
visualizer("<img>What is shown in this image?")
Refer to the Training a better GPT model: Learnings from PaLM blog post for more details about how ABEJA trained GPT-NeoX-Japanese.
( transformers_version: str | None = Nonearchitectures: list[str] | None = Noneoutput_hidden_states: bool | None = Falsereturn_dict: bool | None = Truedtype: typing.Union[str, ForwardRef('torch.dtype'), NoneType] = Nonechunk_size_feed_forward: int = 0is_encoder_decoder: bool = Falseid2label: dict[int, str] | dict[str, str] | None = Nonelabel2id: dict[str, int] | dict[str, str] | None = Noneproblem_type: typing.Optional[typing.Literal['regression', 'single_label_classification', 'multi_label_classification']] = Nonevocab_size: int = 32000hidden_size: int = 2560num_hidden_layers: int = 32num_attention_heads: int = 32intermediate_multiple_size: int = 4hidden_act: str = 'gelu'max_position_embeddings: int = 2048initializer_range: float = 0.02layer_norm_eps: float = 1e-05use_cache: bool = Truebos_token_id: int | None = 31996eos_token_id: int | list[int] | None = 31999rope_parameters: transformers.modeling_rope_utils.RopeParameters | dict | None = Noneattention_dropout: float | int = 0.1hidden_dropout: float | int = 0.0is_decoder: bool = Falsepad_token_id: int | None = Nonetie_word_embeddings: bool = True )
Parameters
int, optional, defaults to 32000) —
Vocabulary size of the model. Defines the number of different tokens that can be represented by the input_ids.int, optional, defaults to 2560) —
Dimension of the hidden representations.int, optional, defaults to 32) —
Number of hidden layers in the Transformer decoder.int, optional, defaults to 32) —
Number of attention heads for each attention layer in the Transformer decoder.int, optional, defaults to 4) —
Dimension of the “intermediate” layer in the Transformer encoder is calculated by hidden_size *
intermediate_multiple_size.str, optional, defaults to gelu) —
The non-linear activation function (function or string) in the decoder. For example, "gelu",
"relu", "silu", etc.int, optional, defaults to 2048) —
The maximum sequence length that this model might ever be used with.float, optional, defaults to 0.02) —
The standard deviation of the truncated_normal_initializer for initializing all weight matrices.float, optional, defaults to 1e-05) —
The epsilon used by the layer normalization layers.bool, optional, defaults to True) —
Whether or not the model should return the last key/values attentions (not used by all models). Only
relevant if config.is_decoder=True or when the model is a decoder-only generative model.int, optional, defaults to 31996) —
Token id used for beginning-of-stream in the vocabulary.Union[int, list[int]], optional, defaults to 31999) —
Token id used for end-of-stream in the vocabulary.Union[~modeling_rope_utils.RopeParameters, dict], optional) —
Dictionary containing the configuration parameters for the RoPE embeddings. The dictionary should contain
a value for rope_theta and optionally parameters used for scaling in case you want to use RoPE
with longer max_position_embeddings.Union[float, int], optional, defaults to 0.1) —
The dropout ratio for the attention probabilities.Union[float, int], optional, defaults to 0.0) —
The dropout probability for all fully connected layers in the embeddings, encoder, and pooler.bool, optional, defaults to False) —
Whether the model is used as a decoder or not. If False, the model is used as an encoder.int, optional) —
Token id used for padding in the vocabulary.bool, optional, defaults to True) —
Whether to tie weight embeddings according to model’s tied_weights_keys mapping.This is the configuration class to store the configuration of a GPTNeoXJapaneseModel. It is used to instantiate a Gpt Neox Japanese model according to the specified arguments, defining the model architecture. Instantiating a configuration with the defaults will yield a similar configuration to that of the abeja/gpt-neox-japanese-2.7b
Configuration objects inherit from PreTrainedConfig and can be used to control the model outputs. Read the documentation from PreTrainedConfig for more information.
Example:
>>> from transformers import GPTNeoXJapaneseConfig, GPTNeoXJapaneseModel
>>> # Initializing a GPTNeoXJapanese gpt-neox-japanese-2.7b style configuration
>>> configuration = GPTNeoXJapaneseConfig()
>>> # Initializing a model (with random weights) from the gpt-neox-japanese-2.7b style configuration
>>> model = GPTNeoXJapaneseModel(configuration)
>>> # Accessing the model configuration
>>> configuration = model.config( vocab_fileemoji_fileunk_token = '<|endoftext|>'pad_token = '<|endoftext|>'bos_token = '<|startoftext|>'eos_token = '<|endoftext|>'do_clean_text = False**kwargs )
Parameters
str) —
File containing the vocabulary.str) —
File containing the emoji.str, optional, defaults to "<|endoftext|>") —
The unknown token. A token that is not in the vocabulary cannot be converted to an ID and is set to be this
token instead.str, optional, defaults to "<|endoftext|>") —
The token used for paddingstr, optional, defaults to "<|startoftext|>") —
The beginning of sequence token.str, optional, defaults to "<|endoftext|>") —
The end of sequence token.bool, optional, defaults to False) —
Whether or not to clean text for URL, EMAIL, TEL, Japanese DATE and Japanese PRICE.This tokenizer inherits from PreTrainedTokenizer and is based on Japanese special Sub-Word-Encoding that is used in this repository (https://github.com/tanreinama/Japanese-BPEEncoder_V2). Check the repository for details. Japanese has a relatively large vocabulary and there is no separation between words. Furthermore, the language is a combination of hiragana, katakana, and kanji, and variants such as “1” and “①” are often used. In order to cope with these, this tokenizer has the following features
Example:
>>> from transformers import GPTNeoXJapaneseTokenizer
>>> tokenizer = GPTNeoXJapaneseTokenizer.from_pretrained("abeja/gpt-neox-japanese-2.7b")
>>> # You can confirm both 慶応 and 慶應 are encoded to 17749
>>> tokenizer("吾輩は猫である🐯。実は慶応(慶應)大学出身")["input_ids"]
[30014, 26883, 26638, 27228, 25, 26650, 31732, 31679, 27809, 26638, 17749, 31592, 17749, 31593, 321, 1281]
>>> # Both 慶応 and 慶應 are decoded to 慶応
>>> tokenizer.decode(tokenizer("吾輩は猫である🐯。実は慶応(慶應)大学出身")["input_ids"])
'吾輩は猫である🐯。実は慶応(慶応)大学出身'Converts a sequence of tokens (string) in a single string.
( config )
Parameters
The bare Gpt Neox Japanese Model outputting raw hidden-states without any specific head on top.
This model inherits from PreTrainedModel. Check the superclass documentation for the generic methods the library implements for all its model (such as downloading or saving, resizing the input embeddings, pruning heads etc.)
This model is also a PyTorch torch.nn.Module subclass. Use it as a regular PyTorch Module and refer to the PyTorch documentation for all matter related to general usage and behavior.
( input_ids: typing.Optional[torch.LongTensor] = Noneattention_mask: typing.Optional[torch.FloatTensor] = Noneposition_ids: typing.Optional[torch.LongTensor] = Noneinputs_embeds: typing.Optional[torch.FloatTensor] = Nonepast_key_values: transformers.cache_utils.Cache | None = Noneuse_cache: bool | None = Noneoutput_attentions: bool | None = Noneoutput_hidden_states: bool | None = Nonereturn_dict: bool | None = None**kwargs ) → BaseModelOutputWithPast or tuple(torch.FloatTensor)
Parameters
input_ids (torch.LongTensor of shape (batch_size, sequence_length), optional) —
Indices of input sequence tokens in the vocabulary. Padding will be ignored by default.
Indices can be obtained using AutoTokenizer. See PreTrainedTokenizer.encode() and PreTrainedTokenizer.call() for details.
What are input IDs? attention_mask (torch.FloatTensor of shape (batch_size, sequence_length), optional) —
Mask to avoid performing attention on padding token indices. Mask values selected in [0, 1]:
torch.LongTensor of shape (batch_size, sequence_length), optional) —
Indices of positions of each input sequence tokens in the position embeddings. Selected in the range [0, config.n_positions - 1].
What are position IDs?torch.FloatTensor of shape (batch_size, sequence_length, hidden_size), optional) —
Optionally, instead of passing input_ids you can choose to directly pass an embedded representation. This
is useful if you want more control over how to convert input_ids indices into associated vectors than the
model’s internal embedding lookup matrix.~cache_utils.Cache, optional) —
Pre-computed hidden-states (key and values in the self-attention blocks and in the cross-attention
blocks) that can be used to speed up sequential decoding. This typically consists in the past_key_values
returned by the model at a previous stage of decoding, when use_cache=True or config.use_cache=True.
Only Cache instance is allowed as input, see our kv cache guide.
If no past_key_values are passed, DynamicCache will be initialized by default.
The model will output the same cache format that is fed as input.
If past_key_values are used, the user is expected to input only unprocessed input_ids (those that don’t
have their past key value states given to this model) of shape (batch_size, unprocessed_length) instead of all input_ids
of shape (batch_size, sequence_length).
bool, optional) —
If set to True, past_key_values key value states are returned and can be used to speed up decoding (see
past_key_values).bool, optional) —
Whether or not to return the attentions tensors of all attention layers. See attentions under returned
tensors for more detail.bool, optional) —
Whether or not to return the hidden states of all layers. See hidden_states under returned tensors for
more detail.bool, optional) —
Whether or not to return a ModelOutput instead of a plain tuple.Returns
BaseModelOutputWithPast or tuple(torch.FloatTensor)
A BaseModelOutputWithPast or a tuple of
torch.FloatTensor (if return_dict=False is passed or when config.return_dict=False) comprising various
elements depending on the configuration (GPTNeoXJapaneseConfig) and inputs.
The GPTNeoXJapaneseModel forward method, overrides the __call__ special method.
Although the recipe for forward pass needs to be defined within this function, one should call the
Moduleinstance afterwards instead of this since the former takes care of running the pre and post processing steps while the latter silently ignores them.
last_hidden_state (torch.FloatTensor of shape (batch_size, sequence_length, hidden_size)) — Sequence of hidden-states at the output of the last layer of the model.
If past_key_values is used only the last hidden-state of the sequences of shape (batch_size, 1, hidden_size) is output.
past_key_values (Cache, optional, returned when use_cache=True is passed or when config.use_cache=True) — It is a Cache instance. For more details, see our kv cache guide.
Contains pre-computed hidden-states (key and values in the self-attention blocks and optionally if config.is_encoder_decoder=True in the cross-attention blocks) that can be used (see past_key_values input) to speed up sequential decoding.
hidden_states (tuple(torch.FloatTensor), optional, returned when output_hidden_states=True is passed or when config.output_hidden_states=True) — Tuple of torch.FloatTensor (one for the output of the embeddings, if the model has an embedding layer, +
one for the output of each layer) of shape (batch_size, sequence_length, hidden_size).
Hidden-states of the model at the output of each layer plus the optional initial embedding outputs.
attentions (tuple(torch.FloatTensor), optional, returned when output_attentions=True is passed or when config.output_attentions=True) — Tuple of torch.FloatTensor (one for each layer) of shape (batch_size, num_heads, sequence_length, sequence_length).
Attentions weights after the attention softmax, used to compute the weighted average in the self-attention heads.
Example:
>>> from transformers import AutoTokenizer, GPTNeoXJapaneseModel
>>> import torch
>>> tokenizer = AutoTokenizer.from_pretrained("abeja/gpt-neox-japanese-2.7b")
>>> model = GPTNeoXJapaneseModel.from_pretrained("abeja/gpt-neox-japanese-2.7b")
>>> inputs = tokenizer("日本語のGPT-neoxがHugging Faceで使えます😀", return_tensors="pt")
>>> outputs = model(**inputs)
>>> last_hidden_states = outputs.last_hidden_state( config )
Parameters
GPTNeoXJapanese Model with a language modeling head on top for Classifier Model fine-tuning.
This model inherits from PreTrainedModel. Check the superclass documentation for the generic methods the library implements for all its model (such as downloading or saving, resizing the input embeddings, pruning heads etc.)
This model is also a PyTorch torch.nn.Module subclass. Use it as a regular PyTorch Module and refer to the PyTorch documentation for all matter related to general usage and behavior.
( input_ids: typing.Optional[torch.LongTensor] = Noneattention_mask: typing.Optional[torch.FloatTensor] = Noneposition_ids: typing.Optional[torch.LongTensor] = Noneinputs_embeds: typing.Optional[torch.FloatTensor] = Nonepast_key_values: transformers.cache_utils.Cache | None = Nonelabels: typing.Optional[torch.LongTensor] = Noneuse_cache: bool | None = Noneoutput_attentions: bool | None = Noneoutput_hidden_states: bool | None = Nonereturn_dict: bool | None = Nonelogits_to_keep: typing.Union[int, torch.Tensor] = 0**kwargs ) → CausalLMOutputWithPast or tuple(torch.FloatTensor)
Parameters
input_ids (torch.LongTensor of shape (batch_size, sequence_length), optional) —
Indices of input sequence tokens in the vocabulary. Padding will be ignored by default.
Indices can be obtained using AutoTokenizer. See PreTrainedTokenizer.encode() and PreTrainedTokenizer.call() for details.
What are input IDs? attention_mask (torch.FloatTensor of shape (batch_size, sequence_length), optional) —
Mask to avoid performing attention on padding token indices. Mask values selected in [0, 1]:
torch.LongTensor of shape (batch_size, sequence_length), optional) —
Indices of positions of each input sequence tokens in the position embeddings. Selected in the range [0, config.n_positions - 1].
What are position IDs?torch.FloatTensor of shape (batch_size, sequence_length, hidden_size), optional) —
Optionally, instead of passing input_ids you can choose to directly pass an embedded representation. This
is useful if you want more control over how to convert input_ids indices into associated vectors than the
model’s internal embedding lookup matrix.~cache_utils.Cache, optional) —
Pre-computed hidden-states (key and values in the self-attention blocks and in the cross-attention
blocks) that can be used to speed up sequential decoding. This typically consists in the past_key_values
returned by the model at a previous stage of decoding, when use_cache=True or config.use_cache=True.
Only Cache instance is allowed as input, see our kv cache guide.
If no past_key_values are passed, DynamicCache will be initialized by default.
The model will output the same cache format that is fed as input.
If past_key_values are used, the user is expected to input only unprocessed input_ids (those that don’t
have their past key value states given to this model) of shape (batch_size, unprocessed_length) instead of all input_ids
of shape (batch_size, sequence_length).
torch.LongTensor of shape (batch_size, sequence_length), optional) —
Labels for computing the left-to-right language modeling loss (next word prediction). Indices should be in
[-100, 0, ..., config.vocab_size] (see input_ids docstring) Tokens with indices set to -100 are
ignored (masked), the loss is only computed for the tokens with labels n [0, ..., config.vocab_size].bool, optional) —
If set to True, past_key_values key value states are returned and can be used to speed up decoding (see
past_key_values).bool, optional) —
Whether or not to return the attentions tensors of all attention layers. See attentions under returned
tensors for more detail.bool, optional) —
Whether or not to return the hidden states of all layers. See hidden_states under returned tensors for
more detail.bool, optional) —
Whether or not to return a ModelOutput instead of a plain tuple.Union[int, torch.Tensor], optional, defaults to 0) —
If an int, compute logits for the last logits_to_keep tokens. If 0, calculate logits for all
input_ids (special case). Only last token logits are needed for generation, and calculating them only for that
token can save memory, which becomes pretty significant for long sequences or large vocabulary size.
If a torch.Tensor, must be 1D corresponding to the indices to keep in the sequence length dimension.
This is useful when using packed tensor format (single dimension for batch and sequence length).Returns
CausalLMOutputWithPast or tuple(torch.FloatTensor)
A CausalLMOutputWithPast or a tuple of
torch.FloatTensor (if return_dict=False is passed or when config.return_dict=False) comprising various
elements depending on the configuration (GPTNeoXJapaneseConfig) and inputs.
The GPTNeoXJapaneseForCausalLM forward method, overrides the __call__ special method.
Although the recipe for forward pass needs to be defined within this function, one should call the
Moduleinstance afterwards instead of this since the former takes care of running the pre and post processing steps while the latter silently ignores them.
loss (torch.FloatTensor of shape (1,), optional, returned when labels is provided) — Language modeling loss (for next-token prediction).
logits (torch.FloatTensor of shape (batch_size, sequence_length, config.vocab_size)) — Prediction scores of the language modeling head (scores for each vocabulary token before SoftMax).
past_key_values (Cache, optional, returned when use_cache=True is passed or when config.use_cache=True) — It is a Cache instance. For more details, see our kv cache guide.
Contains pre-computed hidden-states (key and values in the self-attention blocks) that can be used (see past_key_values input) to speed up sequential decoding.
hidden_states (tuple(torch.FloatTensor), optional, returned when output_hidden_states=True is passed or when config.output_hidden_states=True) — Tuple of torch.FloatTensor (one for the output of the embeddings, if the model has an embedding layer, +
one for the output of each layer) of shape (batch_size, sequence_length, hidden_size).
Hidden-states of the model at the output of each layer plus the optional initial embedding outputs.
attentions (tuple(torch.FloatTensor), optional, returned when output_attentions=True is passed or when config.output_attentions=True) — Tuple of torch.FloatTensor (one for each layer) of shape (batch_size, num_heads, sequence_length, sequence_length).
Attentions weights after the attention softmax, used to compute the weighted average in the self-attention heads.
Example:
>>> from transformers import AutoTokenizer, GPTNeoXJapaneseForCausalLM, GPTNeoXJapaneseConfig
>>> import torch
>>> tokenizer = AutoTokenizer.from_pretrained("abeja/gpt-neox-japanese-2.7b")
>>> config = GPTNeoXJapaneseConfig.from_pretrained("abeja/gpt-neox-japanese-2.7b")
>>> config.is_decoder = True
>>> model = GPTNeoXJapaneseForCausalLM.from_pretrained("abeja/gpt-neox-japanese-2.7b", config=config)
>>> inputs = tokenizer("日本語のGPT-neoxがHugging Faceで使えます😀", return_tensors="pt")
>>> outputs = model(**inputs)
>>> prediction_logits = outputs.logits