January 2018
Volume 33 Number 1
[C#]
All About Span: Exploring a New .NET Mainstay
By Stephen Toub | January 2018
This article has been updated to reflect technical changes following its publication.
Imagine you’re exposing a specialized sort routine to operate in-place on data in memory. You’d likely expose a method that takes an array and provide an implementation that operates over that T[]. That’s great if your method’s caller has an array and wants the whole array sorted, but what if the caller only wants part of it sorted? You’d probably then also expose an overload that took an offset and a count. But what if you wanted to support data in memory that wasn’t in an array, but instead came from native code, for example, or lived on the stack and you only had a pointer and a length? How could you write your sort method that operated on such an arbitrary region of memory, and yet worked equally well with full arrays or with subsets of arrays, and that also worked equally well with managed arrays and unmanaged pointers?
Or take another example. You’re implementing an operation over System.String, such as a specialized parsing method. You’d likely expose a method that takes a string and provide an implementation that operates on strings. But what if you wanted to support operating over a subset of that string? String.Substring could be used to carve out just the piece that’s interesting to them, but that’s a relatively expensive operation, involving a string allocation and memory copy. You could, as mentioned in the array example, take an offset and a count, but then what if the caller doesn’t have a string but instead has a char[]? Or what if the caller has a char*, like one they created with stackalloc to use some space on the stack, or as the result of a call to native code? How could you write your parsing method in a way that didn’t force the caller to do any allocations or copies, and yet worked equally well with inputs of type string, char[] and char*?
In both situations, you might be able to use unsafe code and pointers, exposing an implementation that accepted a pointer and a length. That, however, eliminates the safety guarantees that are core to .NET and opens you up to problems like buffer overruns and access violations that for most .NET developers are a thing of the past. It also invites additional performance penalties, such as needing to pin managed objects for the duration of the operation so that the pointer you retrieve remains valid. And depending on the type of data involved, getting a pointer at all may not be practical.
There’s an answer to this conundrum, and its name is Span
What Is Span?
System.Span
For example, you can create a Span
var arr = new byte[10];
Span bytes = arr; // Implicit cast from T[] to Span
From there, you can easily and efficiently create a span to represent/point to just a subset of this array, utilizing an overload of the span’s Slice method. From there you can index into the resulting span to write and read data in the relevant portion of the original array:
Span slicedBytes = bytes.Slice(start: 5, length: 2);
slicedBytes[0] = 42;
slicedBytes[1] = 43;
Assert.Equal(42, slicedBytes[0]);
Assert.Equal(43, slicedBytes[1]);
Assert.Equal(arr[5], slicedBytes[0]);
Assert.Equal(arr[6], slicedBytes[1]);
slicedBytes[2] = 44; // Throws IndexOutOfRangeException
bytes[2] = 45; // OK
Assert.Equal(arr[2], bytes[2]);
Assert.Equal(45, arr[2]);
As mentioned, spans are more than just a way to access and subset arrays. They can also be used to refer to data on the stack. For example:
Span bytes = stackalloc byte[2]; // Using C# 7.2 stackalloc support for spans
bytes[0] = 42;
bytes[1] = 43;
Assert.Equal(42, bytes[0]);
Assert.Equal(43, bytes[1]);
bytes[2] = 44; // throws IndexOutOfRangeException
More generally, they can be used to refer to arbitrary pointers and lengths, such as to memory allocated from a native heap, like so:
IntPtr ptr = Marshal.AllocHGlobal(1);
try
{
Span bytes;
unsafe { bytes = new Span((byte*)ptr, 1); }
bytes[0] = 42;
Assert.Equal(42, bytes[0]);
Assert.Equal(Marshal.ReadByte(ptr), bytes[0]);
bytes[1] = 43; // Throws IndexOutOfRangeException
}
finally { Marshal.FreeHGlobal(ptr); }
The Span
public ref T this[int index] { get { ... } }
The impact of this ref-returning indexer is most obvious via example, such as by comparing it with the List
struct MutableStruct { public int Value; }
...
Span spanOfStructs = new MutableStruct[1];
spanOfStructs[0].Value = 42;
Assert.Equal(42, spanOfStructs[0].Value);
var listOfStructs = new List { new MutableStruct() };
listOfStructs[0].Value = 42; // Error CS1612: the return value is not a variable
A second variant of Span
string str = "hello, world";
string worldString = str.Substring(startIndex: 7, length: 5); // Allocates
ReadOnlySpan worldSpan =
str.AsSpan().Slice(start: 7, length: 5); // No allocation
Assert.Equal('w', worldSpan[0]);
worldSpan[0] = 'a'; // Error CS0200: indexer cannot be assigned to
Spans provide a multitude of benefits beyond those already mentioned. For example, spans support the notion of reinterpret casts, meaning you can cast a Span
How Is Span Implemented?
Developers generally don’t need to understand how a library they’re using is implemented. However, in the case of Span
First, Span
public readonly ref struct Span
{
private readonly ref T _pointer;
private readonly int _length;
...
}
The concept of a ref T field may be strange at first—in fact, one can’t actually declare a ref T field in C# or even in MSIL. But Span
public static void AddOne(ref int value) => value += 1;
...
var values = new int[] { 42, 84, 126 };
AddOne(ref values[2]);
Assert.Equal(127, values[2]);
This code passes a slot in the array by reference, such that (optimizations aside) you have a ref T on the stack. The ref T in the Span
From this brief description, two things should be clear:
- Span
is defined in such a way that operations can be as efficient as on arrays: indexing into a span doesn’t require computation to determine the beginning from a pointer and its starting offset, as the ref field itself already encapsulates both. (By contrast, ArraySegment has a separate offset field, making it more expensive both to index into and to pass around.) - The nature of Span
as a ref-like type brings with it some constraints due to its ref T field.
This second item has some interesting ramifications that result in .NET containing a second and related set of types, led by Memory
What Is Memory and Why Do You Need It?
Span
var arr = new byte[100];
Span interiorRef1 = arr.AsSpan(start: 20);
Span interiorRef2 = new Span(arr, 20, arr.Length – 20);
Span interiorRef3 =
MemoryMarshal.CreateSpan(arr, ref arr[20], arr.Length – 20);
These references are called interior pointers, and tracking them is a relatively expensive operation for the .NET runtime’s garbage collector. As such, the runtime constrains these refs to only live on the stack, as it provides an implicit low limit on the number of interior pointers that might be in existence.
Further, Span
As a result, Span
These limitations are immaterial for many scenarios, in particular for compute-bound and synchronous processing functions. But asynchronous functionality is another story. Most of the issues cited at the beginning of this article around arrays, array slices, native memory, and so on exist whether dealing with synchronous or asynchronous operations. Yet, if Span
Memory
public readonly struct Memory
{
private readonly object _object;
private readonly int _index;
private readonly int _length;
...
}
You can create a Memory
static async Task ChecksumReadAsync(Memory buffer, Stream stream)
{
int bytesRead = await stream.ReadAsync(buffer);
return Checksum(buffer.Span.Slice(0, bytesRead));
// Or buffer.Slice(0, bytesRead).Span
}
static int Checksum(Span buffer) { ... }
As with Span
Figure 1 Non-Allocating/Non-Copying Conversions Between Span-Related Types
| From | To | Mechanism |
| ArraySegment |
Memory |
Implicit cast, AsMemory method |
| ArraySegment |
ReadOnlyMemory |
Implicit cast, AsMemory method |
| ArraySegment |
ReadOnlySpan |
Implicit cast, AsSpan method |
| ArraySegment |
Span |
Implicit cast, AsSpan method |
| ArraySegment |
T[] | Array property |
| Memory |
ArraySegment |
MemoryMarshal.TryGetArray method |
| Memory |
ReadOnlyMemory |
Implicit cast, AsMemory method |
| Memory |
Span |
Span property |
| ReadOnlyMemory |
ArraySegment |
MemoryMarshal.TryGetArray method |
| ReadOnlyMemory |
ReadOnlySpan |
Span property |
| ReadOnlySpan |
ref readonly T | Indexer get accessor, marshaling methods |
| Span |
ReadOnlySpan |
Implicit cast, AsSpan method |
| Span |
ref T | Indexer get accessor, marshaling methods |
| String | ReadOnlyMemory |
AsMemory method |
| String | ReadOnlySpan |
Implicit cast, AsSpan method |
| T[] | ArraySegment |
Ctor, Implicit cast |
| T[] | Memory |
Ctor, Implicit cast, AsMemory method |
| T[] | ReadOnlyMemory |
Ctor, Implicit cast, AsMemory method |
| T[] | ReadOnlySpan |
Ctor, Implicit cast, AsSpan method |
| T[] | Span |
Ctor, Implicit cast, AsSpan method |
| void* | ReadOnlySpan |
Ctor |
| void* | Span |
Ctor |
You’ll notice that Memory
How Do Span and Memory Integrate with .NET Libraries?
In the previous Memory
In support of Span
string input = ...;
int commaPos = input.IndexOf(',');
int first = int.Parse(input.Substring(0, commaPos));
int second = int.Parse(input.Substring(commaPos + 1));
That, however, incurs two string allocations. If you’re writing performance-sensitive code, that may be two string allocations too many. Instead, you can now write this:
string input = ...;
ReadOnlySpan inputSpan = input;
int commaPos = input.IndexOf(',');
int first = int.Parse(inputSpan.Slice(0, commaPos));
int second = int.Parse(inputSpan.Slice(commaPos + 1));
By using the new Span-based Parse overloads, you’ve made this whole operation allocation-free. Similar parsing and formatting methods exist for primitives like Int32 up through core types like DateTime, TimeSpan and Guid, and even up to higher-level types like BigInteger and IPAddress.
In fact, many such methods have been added across the framework. From System.Random to System.Text.StringBuilder to System.Net.Sockets, overloads have been added to make working with {ReadOnly}Span
public virtual ValueTask ReadAsync(
Memory destination,
CancellationToken cancellationToken = default) { ... }
You’ll notice that unlike the existing ReadAsync method that accepts a byte[] and returns a Task
Because it’s quite common for Stream implementations to buffer in a way that makes ReadAsync calls complete synchronously, this new ReadAsync overload returns a ValueTask
In addition, there are places where Span
int length = ...;
Random rand = ...;
var chars = new char[length];
for (int i = 0; i < chars.Length; i++)
{
chars[i] = (char)(rand.Next(0, 10) + '0');
}
string id = new string(chars);
You could instead use stack-allocation, and even take advantage of Span
int length = ...;
Random rand = ...;
Span chars = stackalloc char[length];
for (int i = 0; i < chars.Length; i++)
{
chars[i] = (char)(rand.Next(0, 10) + '0');
}
string id = new string(chars);
This is better, in that you’ve avoided the heap allocation, but you’re still forced to copy into the string the data that was generated on the stack. This approach also only works when the amount of space required is something small enough for the stack. If the length is short, like 32 bytes, that’s fine, but if it’s thousands of bytes, it could easily lead to a stack overflow situation. What if you could write to the string’s memory directly instead? Span
public static string Create(
int length, TState state, SpanAction action);
...
public delegate void SpanAction(Span span, TArg arg);
This method is implemented to allocate the string and then hand out a writable span you can write to in order to fill in the contents of the string while it’s being constructed. Note that the stack-only nature of Span
int length = ...;
Random rand = ...;
string id = string.Create(length, rand, (Span chars, Random r) =>
{
for (int i = 0; chars.Length; i++)
{
chars[i] = (char)(r.Next(0, 10) + '0');
}
});
Now, not only have you avoided the allocation, you’re writing directly into the string’s memory on the heap, which means you’re also avoiding the copy and you’re not constrained by size limitations of the stack.
Beyond core framework types gaining new members, many new .NET types are being developed to work with spans for efficient processing in specific scenarios. For example, developers looking to write high-performance microservices and Web sites heavy in text processing can earn a significant performance win if they don’t have to encode to and decode from strings when working in UTF-8. To enable this, new types like System.Buffers.Text.Base64, System.Buffers.Text.Utf8Parser and System.Buffers.Text.Utf8Formatter are being added. These operate on spans of bytes, which not only avoids the Unicode encoding and decoding, but enables them to work with native buffers that are common in the very lowest levels of various networking stacks:
ReadOnlySpan utf8Text = ...;
if (!Utf8Parser.TryParse(utf8Text, out Guid value,
out int bytesConsumed, standardFormat = 'P'))
throw new InvalidDataException();
All this functionality isn’t just for public consumption; rather the framework itself is able to utilize these new Span
This doesn’t stop at the level of the core .NET libraries; it continues all the way up the stack. ASP.NET Core now has a heavy dependency on spans, for example, with the Kestrel server’s HTTP parser written on top of them. In the future, it’s likely that spans will be exposed out of public APIs in the lower levels of ASP.NET Core, such as in its middleware pipeline.
What About the .NET Runtime?
One of the ways the .NET runtime provides safety is by ensuring that indexing into an array doesn’t allow going beyond the length of the array, a practice known as bounds checking. For example, consider this method:
[MethodImpl(MethodImplOptions.NoInlining)]
static int Return4th(int[] data) => data[3];
On the x64 machine on which I'm typing this article, the generated assembly for this method looks like the following:
sub rsp, 40
cmp dword ptr [rcx+8], 3
jbe SHORT G_M22714_IG04
mov eax, dword ptr [rcx+28]
add rsp, 40
ret
G_M22714_IG04:
call CORINFO_HELP_RNGCHKFAIL
int3
That cmp instruction is comparing the length of the data array against the index 3, and the subsequent jbe instruction is then jumping to the range check failure routine if 3 is out of range (for an exception to be thrown). The JIT needs to generate code that ensures such accesses don’t go outside the bounds of the array, but that doesn’t mean that every individual array access needs a bound check. Consider this Sum method:
static int Sum(int[] data)
{
int sum = 0;
for (int i = 0; i < data.Length; i++) sum += data[i];
return sum;
}
The JIT needs to generate code here that ensures the accesses to data[i] don’t go outside the bounds of the array, but because the JIT can tell from the structure of the loop that i will always be in range (the loop iterates through each element from beginning to end), the JIT can optimize away the bounds checks on the array. Thus, the assembly code generated for the loop looks like the following:
G_M33811_IG03:
movsxd r9, edx
add eax, dword ptr [rcx+4*r9+16]
inc edx
cmp r8d, edx
jg SHORT G_M33811_IG03
A cmp instruction is still in the loop, but simply to compare the value of i (as stored in the edx register) against the length of the array (as stored in the r8d register); no additional bounds checking.
The runtime applies similar optimizations to span (both Span
static int Sum(Span data)
{
int sum = 0;
for (int i = 0; i < data.Length; i++) sum += data[i];
return sum;
}
The generated assembly for this code is almost identical:
G_M33812_IG03:
movsxd r9, r8d
add ecx, dword ptr [rax+4*r9]
inc r8d
cmp r8d, edx
jl SHORT G_M33812_IG03
The assembly code is so similar in part because of the elimination of bounds checks. But also relevant is the JIT’s recognition of the span indexer as an intrinsic, meaning that the JIT generates special code for the indexer, rather than translating its actual IL code into assembly.
All of this is to illustrate that the runtime can apply for spans the same kinds of optimizations it does for arrays, making spans an efficient mechanism for accessing data. More details are available in the blog post at bit.ly/2zywvyI.
What About the C# Language and Compiler?
I’ve already alluded to features added to the C# language and compiler to help make Span
Ref structs. As noted earlier, Span
public ref struct Enumerator
{
private readonly Span _span;
private int _index;
...
}
Stackalloc initialization of spans. In previous versions of C#, the result of stackalloc could only be stored into a pointer local variable. As of C# 7.2, stackalloc can now be used as part of an expression and can target a span, and that can be done without using the unsafe keyword. Thus, instead of writing:
Span bytes;
unsafe
{
byte* tmp = stackalloc byte[length];
bytes = new Span(tmp, length);
}
You can write simply:
Span bytes = stackalloc byte[length];
This is also extremely useful in situations where you need some scratch space to perform an operation, but want to avoid allocating heap memory for relatively small sizes. Previously you had two choices:
- Write two completely different code paths, allocating and operating over stack-based memory and over heap-based memory.
- Pin the memory associated with the managed allocation and then delegate to an implementation also used for the stack-based memory and written with pointer manipulation in unsafe code.
Now, the same thing can be accomplished without code duplication, with safe code and with minimal ceremony:
Span bytes = length <= 128 ? stackalloc byte[length] : new byte[length];
... // Code that operates on the Span
Span usage validation. Because spans can refer to data that might be associated with a given stack frame, it can be dangerous to pass spans around in a way that might enable referring to memory that’s no longer valid. For example, imagine a method that tried to do the following:
static Span FormatGuid(Guid guid)
{
Span chars = stackalloc char[100];
bool formatted = guid.TryFormat(chars, out int charsWritten, "d");
Debug.Assert(formatted);
return chars.Slice(0, charsWritten); // Uh oh
}
Here space is being allocated from the stack and then trying to return a reference to that space, but the moment you return, that space will no longer be valid for use. Thankfully the C# compiler detects such invalid usage with ref structs and fails the compilation with an error:
error CS8352: Cannot use local 'chars' in this context because it may expose referenced variables outside of their declaration scope
What’s Next?
The types, methods, runtime optimizations, and other elements discussed here are on track to being included in .NET Core 2.1. After that, I expect them to make their way into the .NET Framework. The core types like Span
Of course, keep in mind that there can and will be breaking changes between the current preview version and what’s actually delivered in a stable release. Such changes will in large part be due to feedback from developers like you as you experiment with the feature set. So please do give it a try, and keep an eye on the github.com/dotnet/coreclr and github.com/dotnet/corefx repositories for ongoing work. You can also find documentation at aka.ms/ref72.
Ultimately, the success of this feature set relies on developers trying it out, providing feedback, and building their own libraries utilizing these types, all with the goal of providing efficient and safe access to memory in modern .NET programs. We look forward to hearing from you about your experiences, and even better, to working with you on GitHub to improve .NET further.
Stephen Toub works on .NET at Microsoft. You can find him on GitHub at github.com/stephentoub.
Thanks to the following technical experts for reviewing this article: Krzysztof Cwalina, Eric Erhardt, Ahson Khan, Jan Kotas, Jared Parsons, Marek Safar, Vladimir Sadov, Joseph Tremoulet, Bill Wagner, Jan Vorlicek, Karel Zikmund