ProtoBuf format
Protocol Buffers (ProtoBuf) is a language-neutral binary format that typically uses a separate .proto file to define the protocol schema. It's more compact than CBOR, because it assigns integer numbers to fields instead of names.
Kotlin serialization uses proto2 semantics, where all fields are explicitly required or optional.
Add dependencies for ProtoBuf
To use ProtoBuf in your project, add the ProtoBuf serialization library dependency to your build file:
Use ProtoBuf for binary serialization
To serialize objects with ProtoBuf, use the ProtoBuf class with the .encodeToByteArray() and the .decodeFromByteArray() functions:
In this example, the output prints string values as readable text and the remaining bytes in hexadecimal form. The same bytes correspond to the following values in ProtoBuf hex notation:
Assign field numbers for ProtoBuf serialization
By default, ProtoBuf assigns field numbers automatically.
To keep your schema stable over time, use the @ProtoNumber annotation to assign field numbers explicitly, without needing a separate .proto file. For example, @ProtoNumber(1) assigns field number 1 to a property, so the Kotlin serialization's ProtoBuf format uses that number during encoding and decoding instead of assigning one automatically.
If you plan to reorder properties, assigning field numbers explicitly keeps the schema stable and aligns with Protobuf's compatibility rules for evolving schemas.
Here's an example:
In this example, the name property uses field number 1, and its encoded tag is 0A. The language property uses field number 3, and its encoded tag is 1A.
In ProtoBuf hex notation, the output is equivalent to the following:
Specify integer encoding in ProtoBuf
ProtoBuf encodes integer properties using varint encoding by default.
To use a different integer encoding for a property, apply the @ProtoType annotation with a ProtoIntegerType value. This annotation affects Byte, Short, Int, Long, and Char properties.
The ProtoIntegerType enum supports three options:
The
DEFAULToption uses varint encoding (intXX), which is optimized for small non-negative numbers. For example, the value of1is encoded in one byte as01.The
SIGNEDoption uses signed ZigZag encoding (sintXX), making it suitable for small signed integers. For example, it encodes the value of-2in one byte as03.The
FIXEDoption uses fixed-width encoding (fixedXX), which always uses a fixed number of bytes. For example, it encodes the value of3as four bytes03 00 00 00.
The following example shows all three supported options:
Encode numeric collections as packed fields
In ProtoBuf, packed fields store repeated primitive numeric values more efficiently by writing the list as a single length-delimited entry instead of repeating the field tag for every element.
You can use the @ProtoPacked annotation to serialize collection types except maps in this form. Packed encoding applies only to repeated primitive numeric fields, and the annotation is ignored for other element types.
According to the Protobuf encoding specification, parsers accept both packed and unpacked repeated numeric fields, so decoding doesn't depend on whether @ProtoPacked is present.
Here's an example:
Represent oneof fields
A oneof field defines a group of fields where only one value can be set at a time.
In Kotlin serialization, you can represent this structure with a polymorphic type.
Consider this ProtoBuf message definition:
To represent this message in Kotlin:
Create a
sealed interfaceorabstract classto represent the fields inside theoneofdeclaration:@Serializable sealed interface IPhoneTypeDefine a class for the entire message. Add a
nameproperty and annotate it with@ProtoNumber(1). Add aphoneproperty of the polymorphicIPhoneTypeand annotate it with@ProtoOneOf:@Serializable data class Data( @ProtoNumber(1) val name: String, @ProtoOneOf val phone: IPhoneType?, )For each field in the
oneofdeclaration, create a subclass with a single property that corresponds to that field. Each subclass can be a regular class, data class, or value class:@Serializable @JvmInline value class HomePhone( val number: String ) : IPhoneType @Serializable data class WorkPhone( val number: String ) : IPhoneTypeAnnotate each subclass property with
@ProtoNumberusing the field number from theoneofdeclaration:@Serializable @JvmInline value class HomePhone( @ProtoNumber(2) val number: String ) : IPhoneType @Serializable data class WorkPhone( @ProtoNumber(3) val number: String ) : IPhoneType
Here's a more detailed example where oneof is used to store either a home phone or a work phone:
The output shows that ProtoBuf encodes only one field from the oneof declaration in each message:
0a03546f6d1203313233represents"Tom"withhome_phone.0a054a657272791a03373839represents"Jerry"withwork_phone.
You can also define a class without the @ProtoOneOf annotation if you only need to deserialize data.
For example:
This way, you can deserialize oneof fields without using a sealed hierarchy.
However, it doesn't enforce exclusivity between homeNumber and workNumber during serialization. If both fields have values, the serialized output may not match the original oneof schema, and another ProtoBuf parser may keep only the last field it reads.
Generate a ProtoBuf schema
Typically, working with ProtoBuf involves using a .proto file and a code generator to create code for serialization and deserialization. However, with Kotlin serialization, you can use Kotlin classes annotated with @Serializable as the source for the schema, making .proto files optional.
This approach simplifies the process when all the code involved is written in Kotlin, but interoperability with other languages often still requires a .proto schema.
To generate this schema, use the ProtoBufSchemaGenerator. It generates a Proto2-compatible schema from one or more SerialDescriptor instances.
This gives you a .proto schema that you can use with other ProtoBuf tools.
Here's an example that generates a .proto schema from a Kotlin data class:
This code generates the following .proto schema: