← All posts

SaaS Factory (The Series) · Part 3 of 3

SaaS Factory (The Series) — Data Partitioning: Storage Strategies in SaaS (Part 3/3)

In this post, we continue talking about data partitioning architectures: storage strategies in SaaS.

6 min read1,397 wordsSections: 13Images: 7Sep 19, 2022

Keywords

Share
Comment

In this post, we continue talking about data partitioning architectures: storage strategies in SaaS.

Hybrid: The business compromise

For many organizations, choosing a strategy is not as simple as selecting the Silo, Bridge or Pool model. Your tenants and your business will have a significant influence on how you approach selecting a storage strategy. In some cases, a team may identify a small collection of its tenants that require the Silo or Bridge model. Once they've made that decision, suppose they feel they have to implement all storage with that model. That artificially limits their ability to embrace tenants who might be open to a Pool model. In fact, it may add cost or complexity for a tier of tenants that don't require the attributes of the Silo or Bridge model. One possible compromise is to build a solution that fully supports “pooled” storage as its foundation. Then you can create a separate database for the tenants that require a Siloed storage solution. The following figure provides an example of this approach in action.

Here, we have two tenants (Tenant 1 and Tenant 2) that are using a Silo model, while the remaining tenants run on a pooled storage model. This is abstracted away by a data access layer that hides the tenant's underlying storage from developers. Although this may add a level of complexity to your data access layer and management profile, it can also give your business a way to tier its offering so that it represents the best of both worlds.

Data Migration

Data migration is one of those areas that is often left out of the evaluation of SaaS storage models. With SaaS, however, consider how your architectural choices will influence your ability to continuously deploy new features and capabilities. While performance and the overall tenant experience are important to emphasize, it is also essential to consider how your storage solution will accommodate ongoing changes to the underlying representation of your data.

Migration and Multitenancy

Each of the multi-tenant storage models requires its own unique approach to handling data migration. In the silo and bridge models, you can migrate data on a tenant-by-tenant basis. Your organization may find this appealing because it lets you carefully migrate each SaaS tenant without exposing all tenants to the possibility of a migration error. However, this approach can add more complexity to the overall orchestration of your deployment lifecycle. Data migration in the pool model can be both appealing and challenging. Migration in a pool model gives you a single point at which, once migrated, all tenants have successfully transitioned to your new data model. On the other hand, any problem introduced during a pool migration can affect all of your tenants. From the start, you should think about how data migration fits into your overall multi-tenant SaaS strategy. If you build this migration orchestration into your delivery pipeline early, you tend to achieve a greater degree of agility in your release process.

Security considerations

Data security must be a priority for SaaS providers. When adopting a multi-tenant strategy, your organization needs a robust security strategy to ensure that tenant data is effectively protected from unauthorized access. Protecting this data and conveying that your system has employed the appropriate security measures are essential to earning the trust of your SaaS customers. The storage strategies you choose will likely use common security patterns supported by AWS. Encrypting data at rest, for example, is a horizontal strategy that can be applied universally across any of the models. This provides a baseline level of security that ensures that — even if there is unauthorized access to the data — it would be useless without the keys needed to decrypt the information.

Multitenancy in DynamoDB

The nature of how data is defined and managed by DynamoDB adds some new twists to how you approach multi-tenancy. While some storage services align well with traditional data partitioning strategies, DynamoDB has a somewhat less direct mapping to the Silo, Bridge and Pool models. With DynamoDB, you have to consider some additional factors when selecting your multi-tenant strategy. The following sections explore the AWS mechanisms commonly used to realize each of the multi-tenant partitioning schemes in DynamoDB.

Silo Model

If you implement a silo model in DynamoDB, you have to find a way to create a grouping of one or more tables associated with a tenant. The approach must also create a secure, controlled view of those tables to satisfy the security requirements of silo customers, preventing any possibility of cross-tenant data access.

Bridge Model

For Bridge, the only question would be whether you can relax some of the isolation requirements described for the silo model. You can achieve this by eliminating the introduction of any table-level IAM policies. Assuming your tenants don't require full isolation, you could argue that removing the IAM policies may simplify your provisioning scheme. However, even in the Bridge model there is merit in isolation. So while eliminating IAM isolation may be appealing, it's still good SaaS practice to leverage constructs and policies that can restrict cross-tenant access.

Pool Model

Implementing the pool model in DynamoDB requires you to take a step back and consider how the service manages data. As data is stored in DynamoDB, the service must continually assess and partition the data to achieve scale. If your data profile is evenly distributed, you can simply rely on this underlying partitioning scheme to optimize the performance and cost profile of your SaaS tenants.

Let's look at an example of how you might bring this solution to life. First, you need a separate table, which we'll call the “tenant lookup table”, to capture and manage the mapping of tenants to their corresponding DynamoDB partition keys. The following figure shows an example of how you might structure your tenant lookup table.

Multitenancy in RDS

How data is built and represented in RDS is very much an extension of unmanaged relational environments. The basic mechanisms available in MySQL, for example, are also available in RDS. That makes achieving multi-tenancy across all RDS flavors relatively straightforward. The following sections describe the various strategies commonly employed to realize the partitioning models in RDS.

Silo Model

You can achieve the silo pattern on AWS in several ways. However, the most common and simplest approach to achieving isolation is to create database instances for each tenant. Through instances, you can reach a level of separation that typically satisfies customers' compliance needs without the overhead of provisioning entirely separate accounts.

Bridge Model

The bridge model gives you the opportunity to have tenants with different schemas and some flexibility when migrating tenant data. You could, for example, have different tenants running different versions of the product at any given time and gradually migrate schema changes on a tenant-by-tenant basis. The following figure provides an example of one way to implement the Bridge model in RDS. In this diagram, you have a single RDS database instance that contains separate customer tables for Tenant1 and Tenant2.

Another option here would be to introduce the notion of separate databases for each tenant within an instance. The terminology varies for each RDS flavor. Some RDS storage containers refer to this as a database; others label it a schema.

Pool Model

The pool model for RDS relies on traditional relational indexing schemes to partition tenant data. As part of moving all tenant data into a shared infrastructure model, you store tenant data in a single RDS instance and the tenants share common tables. These tables are indexed with a unique tenant identifier that is used to access and manage each tenant's data

Conclusions

Although there is no single strategy that universally fits every environment, it's clear that some models align better with the core principles of delivering the SaaS model. In general, pool-based storage approaches align well with the need for a unified way to manage and operate a multi-tenant environment. Having all your tenants share repositories and representations simplifies and unifies your operational and deployment approach, enabling cross-tenant views of health and performance.

In the upcoming posts, we'll talk about Identity and Access: the secrets of Identity and Access control in SaaS.

SaaS Factory (The Series) — How to Build Software as a Service Solutions on AWS

See you then! =)

Comments

Every comment is moderated before it appears here. Nothing is published automatically.

Loading…