Repository navigation
Data Transformation
I'm not very happy with this being in the library and it may be phase out/removed in future versions. I believe FluentStorage should be focussed on being the best polycloud storage library and it should not go into tangental directions. I feel that encryption is way out of scope, and the way sinks/transformations are implemented causes a lot of headache and frustration because there are too many APIs. Are we saying that ALL polycloud API will work seamlessly and magically encrypt/decrypt the contents across ALL providers? What a monumental task to support. And for what? Data transformation is an endless sea of possibilities, and I don't like how few of them are implemented here, putting additional pressure on an already strained implementation.
Transform sinks is another awesome feature of FluentStorage that works across all the storage providers. Transform sinks allow you to transform data stream for both upload and download to somehow transform the underlying stream of data. Examples of transform sinks would be gzipping data transparently, encrypting it, and so on.
Let's say you would like to gzip all of the files that you upload/download to a storage. You can do that in the following way:
IStore myGzippedStorage = {StorageFactory}
.FromXXX()
.WithGzipCompression();Then use the storage as you would before - all the data is compressed as you write it (with any WriteXXX method) and decompressed as you read it (with any ReadXXX method).
Due to the nature of the transforms, they can change both the underlying data, and stream size, therefore there is an issue with storage providers, as they need to know beforehand the size of the blob you are uploading. The matter becomes more complicated when some implementations need to calculate other statistics of the data before uploading i.e. hash, CRC and so on. Therefore the only reliable way to stream transformed data is to actually perform all of the transofrms, and then upload it. In this implementation, FluentStorage uses in-memory transforms to achieve this, however does it extremely efficiently by using Microsoft.IO.RecyclableMemoryStream package that performs memory pooling and reclaiming for you so that you don't need to worry about software slowdows. You can read more about this technique here.
This also means that today a transform sink can upload a stream only as large as the amount of RAM available on your machine. I am, however, thinking of ways to go further than that, and there are some beta implementations available that might see the light soon.
- AWS S3 Storage
- Azure Blob Storage
- Azure File Storage
- Azure Data Lake
- Azure Key Vault
- GCP Storage
- Cloudflare R2 Storage
- MinIO Storage
- DigitalOcean Spaces
- Wasabi Storage
- Backblaze B2 Storage
- Hetzner Storage
- Vultr Storage
- MongoDB GridFS Storage
- Alibaba OSS Storage
- FTP Storage
- SFTP Storage
- Git Repository Storage