0% found this document useful (0 votes)
9 views19 pages

SVC Class in Scikit-Learn 1.5.1

The document provides detailed documentation for the SVC (Support Vector Classification) class in scikit-learn version 1.5.1, including its parameters, attributes, and methods. It explains the functionality of various parameters such as C, kernel type, and gamma, and outlines how to fit the model and make predictions. Additionally, it includes examples and references for further reading on support vector machines.

Uploaded by

Shayrula Alice
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views19 pages

SVC Class in Scikit-Learn 1.5.1

The document provides detailed documentation for the SVC (Support Vector Classification) class in scikit-learn version 1.5.1, including its parameters, attributes, and methods. It explains the functionality of various parameters such as C, kernel type, and gamma, and outlines how to fit the model and make predictions. Additionally, it includes examples and references for further reading on support vector machines.

Uploaded by

Shayrula Alice
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

18/07/24, 22:13 SVC — scikit-learn 1.5.

1 documentation

Section Navigation SVC


sklearn
class [Link](*, C=1.0, kernel='rbf', degree=3, gamma='scale', coef0=0.0,
[Link]
shrinking=True, probability=False, tol=0.001, cache_size=200, class_weight=None,
[Link] verbose=False, max_iter=-1, decision_function_shape='ovr', break_ties=False,
[Link] random_state=None) [source]
[Link] C-Support Vector Classification.
[Link] The implementation is based on libsvm. The fit time scales at least quadratically with the number of
sklearn.cross_decomposi samples and may be impractical beyond tens of thousands of samples. For large datasets consider using
tion LinearSVC or SGDClassifier instead, possibly after a Nystroem transformer or other Kernel
[Link] Approximation.

[Link] The multiclass support is handled according to a one-vs-one scheme.


sklearn.discriminant_anal For details on the precise mathematical formulation of the provided kernel functions and how gamma ,
ysis coef0 and degree affect each other, see the corresponding section in the narrative documentation:
[Link] Kernel functions.
[Link] To learn how to tune SVC’s hyperparameters, see the following example: Nested versus non-nested
[Link] cross-validation

[Link] Read more in the User Guide.


sklearn.feature_extractio
Parameters:
n
C : float, default=1.0
sklearn.feature_selection
Regularization parameter. The strength of the regularization is inversely proportional to C. Must
be strictly positive. The penalty is a squared l2 penalty. For an intuitive visualization of the effects

[Link] 1/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

of scaling the regularization parameter C, see Scaling the regularization parameter for SVCs.

kernel : {‘linear’, ‘poly’, ‘rbf’, ‘sigmoid’, ‘precomputed’} or callable, default=’rbf’


Specifies the kernel type to be used in the algorithm. If none is given, ‘rbf’ will be used. If a callable
is given it is used to pre-compute the kernel matrix from data matrices; that matrix should be an
array of shape (n_samples, n_samples) . For an intuitive visualization of different kernel types see
Plot classification boundaries with different SVM Kernels.

degree : int, default=3


Degree of the polynomial kernel function (‘poly’). Must be non-negative. Ignored by all other
kernels.

gamma : {‘scale’, ‘auto’} or float, default=’scale’


Kernel coefficient for ‘rbf’, ‘poly’ and ‘sigmoid’.

if gamma='scale' (default) is passed then it uses 1 / (n_features * [Link]()) as value of gamma,

if ‘auto’, uses 1 / n_features

if float, must be non-negative.

 Changed in version 0.22: The default value of gamma changed from ‘auto’ to ‘scale’.

coef0 : float, default=0.0


Independent term in kernel function. It is only significant in ‘poly’ and ‘sigmoid’.

shrinking : bool, default=True


Whether to use the shrinking heuristic. See the User Guide.

probability : bool, default=False


Whether to enable probability estimates. This must be enabled prior to calling fit , will slow
down that method as it internally uses 5-fold cross-validation, and predict_proba may be
[Link] 2/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

inconsistent with predict . Read more in the User Guide.

tol : float, default=1e-3


Tolerance for stopping criterion.

cache_size : float, default=200


Specify the size of the kernel cache (in MB).

class_weight : dict or ‘balanced’, default=None


Set the parameter C of class i to class_weight[i]*C for SVC. If not given, all classes are supposed to
have weight one. The “balanced” mode uses the values of y to automatically adjust weights
inversely proportional to class frequencies in the input data as n_samples / (n_classes *
[Link](y)) .

verbose : bool, default=False


Enable verbose output. Note that this setting takes advantage of a per-process runtime setting in
libsvm that, if enabled, may not work properly in a multithreaded context.

max_iter : int, default=-1


Hard limit on iterations within solver, or -1 for no limit.

decision_function_shape : {‘ovo’, ‘ovr’}, default=’ovr’


Whether to return a one-vs-rest (‘ovr’) decision function of shape (n_samples, n_classes) as all
other classifiers, or the original one-vs-one (‘ovo’) decision function of libsvm which has shape
(n_samples, n_classes * (n_classes - 1) / 2). However, note that internally, one-vs-one (‘ovo’) is
always used as a multi-class strategy to train models; an ovr matrix is only constructed from the
ovo matrix. The parameter is ignored for binary classification.

 Changed in version 0.19: decision_function_shape is ‘ovr’ by default.

[Link] 3/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

 Added in version 0.17: decision_function_shape=’ovr’ is recommended.

 Changed in version 0.17: Deprecated decision_function_shape=’ovo’ and None.

break_ties : bool, default=False


If true, decision_function_shape='ovr' , and number of classes > 2, predict will break ties
according to the confidence values of decision_function; otherwise the first class among the tied
classes is returned. Please note that breaking ties comes at a relatively high computational cost
compared to a simple predict.

 Added in version 0.22.

random_state : int, RandomState instance or None, default=None


Controls the pseudo random number generation for shuffling the data for probability estimates.
Ignored when probability is False. Pass an int for reproducible output across multiple function
calls. See Glossary.

Attributes:

class_weight_ : ndarray of shape (n_classes,)


Multipliers of parameter C for each class. Computed based on the class_weight parameter.

classes_ : ndarray of shape (n_classes,)


The classes labels.

coef_ : ndarray of shape (n_classes * (n_classes - 1) / 2, n_features)

Weights assigned to the features when kernel="linear" .

dual_coef_ : ndarray of shape (n_classes -1, n_SV)

[Link] 4/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Dual coefficients of the support vector in the decision function (see Mathematical formulation),
multiplied by their targets. For multiclass, coefficient for all 1-vs-1 classifiers. The layout of the
coefficients in the multiclass case is somewhat non-trivial. See the multi-class section of the User
Guide for details.

fit_status_ : int
0 if correctly fitted, 1 otherwise (will raise warning)

intercept_ : ndarray of shape (n_classes * (n_classes - 1) / 2,)


Constants in decision function.

n_features_in_ : int
Number of features seen during fit.

 Added in version 0.24.

feature_names_in_ : ndarray of shape ( n_features_in_ ,)


Names of features seen during fit. Defined only when X has feature names that are all strings.

 Added in version 1.0.

n_iter_ : ndarray of shape (n_classes * (n_classes - 1) // 2,)


Number of iterations run by the optimization routine to fit the model. The shape of this attribute
depends on the number of models optimized which in turn depends on the number of classes.

 Added in version 1.1.

support_ : ndarray of shape (n_SV)


Indices of support vectors.
[Link] 5/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

support_vectors_ : ndarray of shape (n_SV, n_features)


Support vectors. An empty array if kernel is precomputed.

n_support_ : ndarray of shape (n_classes,), dtype=int32

Number of support vectors for each class.

probA_ : ndarray of shape (n_classes * (n_classes - 1) / 2)

Parameter learned in Platt scaling when probability=True .

probB_ : ndarray of shape (n_classes * (n_classes - 1) / 2)

Parameter learned in Platt scaling when probability=True .

shape_fit_ : tuple of int of shape (n_dimensions_of_X,)


Array dimensions of training vector X .

 See also

SVR
Support Vector Machine for Regression implemented using libsvm.

LinearSVC
Scalable Linear Support Vector Machine for classification implemented using liblinear. Check
the See Also section of LinearSVC for more comparison element.

References

[1] LIBSVM: A Library for Support Vector Machines

[2] Platt, John (1999). “Probabilistic Outputs for Support Vector Machines and Comparisons to
Regularized Likelihood Methods”
[Link] 6/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Examples

>>> import numpy as np


>>> from [Link] import make_pipeline
>>> from [Link] import StandardScaler
>>> X = [Link]([[-1, -1], [-2, -1], [1, 1], [2, 1]])
>>> y = [Link]([1, 1, 2, 2])
>>> from [Link] import SVC
>>> clf = make_pipeline(StandardScaler(), SVC(gamma='auto'))
>>> [Link](X, y)
Pipeline(steps=[('standardscaler', StandardScaler()),
('svc', SVC(gamma='auto'))])

>>> print([Link]([[-0.8, -1]]))


[1]

property coef_
Weights assigned to the features when kernel="linear" .

Returns:

ndarray of shape (n_features, n_classes)

decision_function(X) [source]

Evaluate the decision function for the samples in X.

Parameters:

X : array-like of shape (n_samples, n_features)


The input samples.

[Link] 7/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Returns:

X : ndarray of shape (n_samples, n_classes * (n_classes-1) / 2)


Returns the decision function of the sample for each class in the model. If
decision_function_shape=’ovr’, the shape is (n_samples, n_classes).

Notes

If decision_function_shape=’ovo’, the function values are proportional to the distance of the samples X
to the separating hyperplane. If the exact distances are required, divide the function values by the
norm of the weight vector ( coef_ ). See also this question for further details. If
decision_function_shape=’ovr’, the decision function is a monotonic transformation of ovo decision
function.

fit(X, y, sample_weight=None) [source]

Fit the SVM model according to the given training data.

Parameters:

X : {array-like, sparse matrix} of shape (n_samples, n_features) or (n_samples, n_samples)


Training vectors, where n_samples is the number of samples and n_features is the number of
features. For kernel=”precomputed”, the expected shape of X is (n_samples, n_samples).

y : array-like of shape (n_samples,)


Target values (class labels in classification, real numbers in regression).

sample_weight : array-like of shape (n_samples,), default=None


Per-sample weights. Rescale C per sample. Higher weights force the classifier to put more
emphasis on these points.

[Link] 8/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Returns:

self : object
Fitted estimator.

Notes

If X and y are not C-ordered and contiguous arrays of np.float64 and X is not a [Link].csr_matrix,
X and/or y may be copied.

If X is a dense array, then the other methods will not support sparse matrices as input.

get_metadata_routing() [source]

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:

routing : MetadataRequest
A MetadataRequest encapsulating routing information.

get_params(deep=True) [source]

Get parameters for this estimator.

Parameters:

deep : bool, default=True


If True, will return the parameters for this estimator and contained subobjects that are
estimators.
[Link] 9/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Returns:

params : dict
Parameter names mapped to their values.

property n_support_
Number of support vectors for each class.

predict(X) [source]

Perform classification on samples in X.

For an one-class model, +1 or -1 is returned.

Parameters:

X : {array-like, sparse matrix} of shape (n_samples, n_features) or (n_samples_test,


n_samples_train)
For kernel=”precomputed”, the expected shape of X is (n_samples_test, n_samples_train).

Returns:

y_pred : ndarray of shape (n_samples,)


Class labels for samples in X.

predict_log_proba(X) [source]

Compute log probabilities of possible outcomes for samples in X.

The model need to have probability information computed at training time: fit with attribute
probability set to True.

[Link] 10/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Parameters:

X : array-like of shape (n_samples, n_features) or (n_samples_test, n_samples_train)


For kernel=”precomputed”, the expected shape of X is (n_samples_test, n_samples_train).

Returns:

T : ndarray of shape (n_samples, n_classes)


Returns the log-probabilities of the sample for each class in the model. The columns
correspond to the classes in sorted order, as they appear in the attribute classes_.

Notes

The probability model is created using cross validation, so the results can be slightly different than
those obtained by predict. Also, it will produce meaningless results on very small datasets.

predict_proba(X) [source]

Compute probabilities of possible outcomes for samples in X.

The model needs to have probability information computed at training time: fit with attribute
probability set to True.

Parameters:

X : array-like of shape (n_samples, n_features)


For kernel=”precomputed”, the expected shape of X is (n_samples_test, n_samples_train).

Returns:

T : ndarray of shape (n_samples, n_classes)

[Link] 11/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Returns the probability of the sample for each class in the model. The columns correspond to
the classes in sorted order, as they appear in the attribute classes_.

Notes

The probability model is created using cross validation, so the results can be slightly different than
those obtained by predict. Also, it will produce meaningless results on very small datasets.

property probA_
Parameter learned in Platt scaling when probability=True .

Returns:

ndarray of shape (n_classes * (n_classes - 1) / 2)

property probB_
Parameter learned in Platt scaling when probability=True .

Returns:

ndarray of shape (n_classes * (n_classes - 1) / 2)

score(X, y, sample_weight=None) [source]

Return the mean accuracy on the given test data and labels.

In multi-label classification, this is the subset accuracy which is a harsh metric since you require for
each sample that each label set be correctly predicted.

Parameters:

[Link] 12/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

X : array-like of shape (n_samples, n_features)


Test samples.

y : array-like of shape (n_samples,) or (n_samples, n_outputs)


True labels for X .

sample_weight : array-like of shape (n_samples,), default=None


Sample weights.

Returns:

score : float
Mean accuracy of [Link](X) w.r.t. y .

set_fit_request(*, sample_weight: bool | None | str = '$UNCHANGED$') → SVC [source]

Request metadata passed to the fit method.

Note that this method is only relevant if enable_metadata_routing=True (see sklearn.set_config ).


Please see User Guide on how the routing mechanism works.

The options for each parameter are:

True : metadata is requested, and passed to fit if provided. The request is ignored if metadata
is not provided.

False : metadata is not requested and the meta-estimator will not pass it to fit .

None : metadata is not requested, and the meta-estimator will raise an error if the user provides it.

str : metadata should be passed to the meta-estimator with this given alias instead of the
original name.

[Link] 13/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

The default ( [Link].metadata_routing.UNCHANGED ) retains the existing request. This allows you
to change the request for some parameters and not others.

 Added in version 1.3.

 Note

This method is only relevant if this estimator is used as a sub-estimator of a meta-estimator,


e.g. used inside a Pipeline . Otherwise it has no effect.

Parameters:

sample_weight : str, True, False, or None,


default=[Link].metadata_routing.UNCHANGED
Metadata routing for sample_weight parameter in fit .

Returns:

self : object
The updated object.

set_params(**params) [source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline ). The latter
have parameters of the form <component>__<parameter> so that it’s possible to update each
component of a nested object.

[Link] 14/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Parameters:

**params : dict
Estimator parameters.

Returns:

self : estimator instance


Estimator instance.

set_score_request(*, sample_weight: bool | None | str = '$UNCHANGED$') → SVC


Request metadata passed to the score method. [source]

Note that this method is only relevant if enable_metadata_routing=True (see sklearn.set_config ).


Please see User Guide on how the routing mechanism works.

The options for each parameter are:

True : metadata is requested, and passed to score if provided. The request is ignored if
metadata is not provided.

False : metadata is not requested and the meta-estimator will not pass it to score .

None : metadata is not requested, and the meta-estimator will raise an error if the user provides it.

str : metadata should be passed to the meta-estimator with this given alias instead of the
original name.

The default ( [Link].metadata_routing.UNCHANGED ) retains the existing request. This allows you
to change the request for some parameters and not others.

 Added in version 1.3.

[Link] 15/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

 Note

This method is only relevant if this estimator is used as a sub-estimator of a meta-estimator,


e.g. used inside a Pipeline . Otherwise it has no effect.

Parameters:

sample_weight : str, True, False, or None,


default=[Link].metadata_routing.UNCHANGED
Metadata routing for sample_weight parameter in score .

Returns:

self : object
The updated object.

Gallery examples

Release Highlights for Release Highlights for Classifier comparison Plot classification
scikit-learn 0.24 scikit-learn 0.22 probability

[Link] 16/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Recognizing hand- Plot the decision Faces recognition Scalable learning with
written digits boundaries of a example using polynomial kernel
VotingClassifier eigenfaces and SVMs approximation

Displaying Pipelines Explicit feature map Multilabel ROC Curve with


approximation for classification Visualization API
RBF kernels

Comparison between Confusion matrix Custom refit strategy Nested versus non-
grid search and of a grid search with nested cross-
successive halving cross-validation validation

[Link] 17/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

Plotting Learning Plotting Validation Receiver Operating Statistical comparison


Curves and Checking Curves Characteristic (ROC) of models using grid
Models' Scalability with cross validation search

Test with Concatenating Feature discretization Decision boundary of


permutations the multiple feature semi-supervised
significance of a extraction methods classifiers versus SVM
classification score on the Iris dataset

Effect of varying Plot classification Plot different SVM RBF SVM parameters
threshold for self- boundaries with classifiers in the iris
training different SVM Kernels dataset

[Link] 18/19
18/07/24, 22:13 SVC — scikit-learn 1.5.1 documentation

SVM Margins SVM Tie Breaking SVM with custom SVM-Anova: SVM
Example Example kernel with univariate
feature selection

SVM: Maximum SVM: Separating SVM: Weighted SVM Exercise

© Copyright 2007 - 2024, scikit-learn developers (BSD License).

[Link] 19/19

Common questions

Powered by AI

Using a precomputed kernel in the SVC class changes the required shape of the input data from the usual (n_samples, n_features) to (n_samples, n_samples). This is because the precomputed kernel approach expects the kernel matrix c, where each entry corresponds to the kernel function computed between a pair of samples, effectively predefining relationships rather than transforming features on-the-fly. While this can save computation time in certain contexts, it requires that the kernel values be generated separately and accurately, and the method does not inherently accommodate additional features or test samples outside those used in precomputation .

The fit time for the SVC class scales quadratically with the number of samples due to the quadratic nature of the optimization problem, where the number of pairwise interactions between samples must be considered, leading to a significant computational burden. For large datasets, this quadratic scaling becomes impractical. As alternatives, LinearSVC or SGDClassifier can be used as they are more suitable for linearly separable support vector machines and can efficiently handle large-scale datasets. Additionally, using kernel approximation techniques, like the Nystroem transformer, can reduce computational cost by approximating the kernel matrix .

The regularization parameter C in the SVC class is inversely proportional to the strength of the regularization. A smaller value of C implies stronger regularization, which helps to avoid overfitting by allowing some misclassifications in exchange for a smaller margin in the decision function. This means that larger values of C correspond to less regularization and aim to classify all training examples correctly by increasing the size of the margin between the decision boundary and data points .

Enabling probability estimates in the SVC class requires additional computational resources as the method internally performs a 5-fold cross-validation to compute the Platt scaling parameters, probA_ and probB_. This added computation can slow down the training process and sometimes yield probability estimates that are not always consistent with the predict method, particularly on very small datasets. Users might enable probability estimates to obtain probabilistic predictions, which can be useful in scenarios requiring class probability ranking or thresholding. However, when computational efficiency is a concern or for datasets where deterministic class assignments suffice, users might opt to disable this feature .

The cache_size parameter in the SVC model specifies the size of the kernel cache in megabytes, which can significantly influence the speed and efficiency of model training. A larger cache size can store more elements of the kernel matrix, reducing the need to recompute them and thus speeding up the fitting process. However, setting the cache size too large may lead to memory issues, particularly on machines with limited RAM. Conversely, a cache size too small will slow down the computation, as elements need to be recalculated frequently. Therefore, optimal cache size should balance these considerations in accordance with available system resources and dataset size .

The computation of the decision function in the SVC model is influenced by the chosen kernel and its parameters like gamma, degree, and coef0, which determine how the input space is transformed. The decision_function_shape parameter impacts the result by controlling the nature of the score matrix; 'ovo' (one-vs-one) provides decision function values for each pair of classes, while 'ovr' (one-vs-rest) provides them for each class, aggregating the votes to yield scores. In 'ovo', scores are derived from pairwise decision boundaries, which can directly reflect the margin distance, whereas 'ovr' averages or combines for final output .

The shrinking heuristic in the SVC model aims to speed up training by temporarily removing certain variables from the optimization when progress stalls. By focusing on a smaller subset of likely support vectors, it reduces the problem size, potentially accelerating convergence. However, this comes at the risk of missing the optimal solution if active variables are re-introduced prematurely or unnecessarily. Users should enable shrinking when dealing with large datasets to benefit from faster computations, but could consider deactivating it for smaller datasets or when every sample might potentially be significant, ensuring all data are considered throughout the optimization process .

The SVC class handles multi-class classification problems using a one-vs-one scheme. In this approach, a binary classifier is trained for every possible pair of classes. Considering n classes, a total of n*(n-1)/2 classifiers are needed. Each classifier distinguishes between two classes, all other classes are ignored. The final decision is made by a voting mechanism among all these classifiers when predicting the class of a sample .

The sample_weight parameter in the SVC model allows users to specify the importance of each sample during training by adjusting the regularization parameter C on a per-sample basis. This is particularly useful in addressing class imbalance, as more weight can be assigned to samples from the minority class, increasing their influence during model training. By applying higher sample_weights to minority samples, the model can focus more effort on these data points, potentially improving classification performance on imbalanced datasets where standard training might result in bias toward the majority class .

The 'rbf', 'poly', and 'sigmoid' kernels in SVM differ as follows: 'rbf' uses the gamma parameter, and its value determines the influence of single training samples, where a lower value of gamma indicates a larger influence and vice versa. The 'poly' kernel is influenced by both the degree and coef0 parameters; it represents the polynomial degree and an independent term in the polynomial kernel. The 'sigmoid' kernel also utilizes the coef0 parameter as an independent term. While 'scale' and 'auto' are options for gamma that affect 'rbf', 'poly', and 'sigmoid', degree is ignored by 'rbf' and 'sigmoid'. Each kernel type transforms the input space differently to find the optimal separator .

You might also like