CBOR format
Concise Binary Object Representation (CBOR) is a compact binary format that extends the JSON data model.
Add dependencies for CBOR
To use CBOR in your project, add the CBOR serialization library dependency to your build file:
Use CBOR for binary serialization
The Cbor class provides two functions:
encodeToByteArray()serializes objects to a byte array.decodeFromByteArray()deserializes objects from a byte array.
Here's an example where a Project object is serialized into a byte array and then deserialized back to its original form:
This example prints the encoded bytes in a readable mixed form. It represents printable ASCII bytes as characters and non-printable bytes as hexadecimal values.
Here's the same output in full CBOR hex notation:
Ignore unknown keys in CBOR
CBOR is commonly used in communication with Internet of Things (IoT) devices where new properties may be added as part of API evolution. By default, unknown keys encountered during deserialization result in an error.
Just like in JSON, you set the ignoreUnknownKeys property to true to ignore them during deserialization:
In this CBOR input, the following bytes represent the unknown "language" key and its value:
The unknown key
"language":68: Length of the key6c616e6775616765: The key
The value
"Kotlin":66: Length of the value4b6f746c696e: The value
Customize CBOR data encoding
According to the RFC 8949 Major Types specification, CBOR supports the following data types:
Major type 0: an unsigned integer
Major type 1: a negative integer
Major type 2: a byte string
Major type 3: a text string
Major type 4: an array of data items
Major type 5: a map of pairs of data items
Major type 6: optional semantic tagging of other major types
Major type 7: floating-point numbers, simple data types with no content, and the "break" stop code
Unlike JSON, CBOR supports maps with structured map keys, such as instances of user-defined classes. However, some parsers, such as jackson-dataformat-cbor, don't support them.
Encode ByteArray properties as byte strings
You can customize the CBOR representation of your data instead of using the default encoding, for example, to match a serialized form defined by an external schema. In some cases, these customizations can also reduce binary size.
By default, Kotlin ByteArray values are encoded as major type 4, which represents an array of data items. To encode ByteArray properties as major type 2, a byte string, use the @ByteString annotation:
In this example, the bytes before each ByteArray value differ because the properties use different CBOR major types.
Here's the encoded byte array in full CBOR hex notation:
To encode all ByteArray values as major type 2 without annotating each property with @ByteString, set the alwaysUseByteString property to true:
Encode classes as CBOR arrays
By default, classes are serialized as a CBOR map, which corresponds to major type 5. This means that each property of the class is stored as a key-value pair.
You can serialize a class as a CBOR array (major type 4) with the @CborArray annotation. This can be useful for encoding COSE message structures, which RFC 9052 defines as CBOR arrays.
Here's an example:
With the @CborArray annotation, this example is encoded as a CBOR array: 0x8226f6. Without it, the same class is encoded as a CBOR map: 0xa263616c6726636b6964f6.
Definite-length and indefinite-length encoding in CBOR
CBOR supports two encodings for maps and arrays: definite-length encoding and indefinite-length encoding.
By default, Kotlin serialization uses indefinite-length encoding. In this encoding, the number of elements in a map or array isn't encoded explicitly. Instead, a "break" stop code (0xFF) marks the end of a collection.
Definite-length encoding omits the terminating byte and encodes the number of elements at the start of the map or array.
To switch between these two modes, use the useDefiniteLengthEncoding property.
Tags and labels in CBOR
CBOR allows you to define tags that encode additional information for properties and values. You can specify these tags with the @KeyTags and @ValueTags annotations. The encodeKeyTags, encodeValueTags, verifyKeyTags, and verifyValueTags properties control the encoding and verification of these tags.
You can also assign tags to all instances of a class with the @ObjectTags annotation.
When serializing, @ObjectTags are encoded directly before the data of the tagged object. If a property has value tags and its type has object tags, the value tags are encoded before the object tags. The encodeObjectTags and verifyObjectTags properties control whether object tags are encoded and verified. If you verify only value tags and don't verify object tags, the decoder can still deserialize data with additional object tags.
CBOR supports map keys of any type. In COSE (CBOR Object Signing and Encryption), these keys are restricted to strings and numbers and are called labels.
You can assign string labels with the @SerialName annotation and numeric labels with the @CborLabel annotation. The preferCborLabelsOverNames property allows prioritizing numeric labels over serial names when both are present. You can use it to keep compact labels for CBOR while still keeping readable names when serializing to JSON.
Kotlin serialization also provides a predefined Cbor.CoseCompliant instance that follows COSE encoding requirements. It uses definite-length encoding, encodes and verifies all tags, and prefers numeric labels over serial names.
Custom CBOR-specific serializers
CBOR encoders and decoders implement the CborEncoder and CborDecoder interfaces.
These interfaces extend the general Encoder and Decoder interfaces, providing access to CBOR-specific configurations through the cbor property. Custom serializers can use this property to access the current Cbor instance, produce embedded byte arrays, and read the current settings, such as preferCborLabelsOverNames and useDefiniteLengthEncoding.
For more information about creating custom serializers, see Create custom serializers.