SaaS Factory (The Series)
SaaS Factory (The Series) — Optimizing SaaS Tenant Costs and Workflows
In this post, we'll talk about optimizing SaaS tenant costs and workflows.
Keywords
In this post, we'll talk about optimizing SaaS tenant costs and workflows.
Let's say, for example, that we're an e-commerce product and we use catalog size as the defining metric that separates the tiers of our system. In this scenario, the business would simply start acquiring customers and onboarding them into the system without worrying about the impact on the bottom line. Now, if we take this system and start collecting per-tenant cost data, the breakdown of infrastructure consumption might end up looking like the bar chart shown below. Here, you'll notice that most tenants signed up for the basic tier. However, the infrastructure costs associated with the basic tier far exceed those of the other two tiers.

Giving the business options
The previous example highlights the close link that often exists between the business and technical sides of a SaaS team. The architectural choices made by SaaS developers have the potential to shape and influence the menu of pricing and packaging options the business offers. The more flexibility you can give the business, the more likely it is to respond quickly to the varied requirements of tenants — each of which may have its own scale, performance and consumption profile.
The broader goal here is to make choices in your SaaS architecture that best align a tenant's expectations with their experience. If a tenant is a US$ 29/month basic tier tenant, they will likely understand that their experience will differ from that of a US$ 5,000/month professional tier tenant. Supporting this model often means introducing variations in policies and, in some cases, architectural models that deliver distinct tenant experiences for each tier of your solution. Examples of how this can be achieved are described in a blog post on how to optimize SaaS tenant workflows and costs, as well as a talk on how to optimize SaaS solutions on AWS .
The hard part: instrumenting your solution
Analyzing Costs in a Silo Model
Some SaaS providers rely exclusively on a siloed partitioning model in which each tenant is housed in practically isolated infrastructure. This isolation may be mandatory for some domains that have strict regulations prohibiting shared infrastructure. For these environments, more pronounced boundaries between each tenant can simplify efforts to aggregate and derive cost analytics.
The other common isolation scheme involves creating separate VPCs for each tenant. With this approach, you can isolate tenants without creating separate accounts for each one. This generally scales better and simplifies the provisioning model. However, it also adds a degree of complexity to cost instrumentation. Instead of relying on linked accounts to summarize tenant spending, you'll need to apply AWS tags to your infrastructure to associate it with your tenants. These tags will then be used to aggregate each tenant's spending.
In a more isolated model, you may want to consider leveraging partner solutions to simplify the aggregation and analysis of your tenant costs. AWS Partner Network (APN) partners such as CloudHealth and Cloudability, for example, provide cost analysis tools that can simplify your ability to perform tenant analysis.
Analyzing costs in a Pool model
In a Pool model, where tenant resources are shared, attributing tenant costs can be more challenging. Here, you must employ more specialized strategies to properly capture and classify tenant activity. You'll need to rely on measurements of tenants' actual interactions with system resources to apportion load and consumption and derive an approximation of tenant costs. This typically requires more effort to instrument your solution and aggregate the resulting metrics.
Let's consider how, in a pooled environment, you might derive a tenant's consumption from its activity. The diagram below depicts a simplified view of an Amazon EC2 Container Service (Amazon ECS) cluster running services in a pooled multi-tenant environment. This cluster includes two services running in containers (Checkout and Catalog), both of which are being called by individual tenants.

As tenants use these services, they consume compute resources. The question is: how much of those resources is actually being consumed by each tenant? Tenant 1, for example, might be pushing the system hard and selling a lot of items, forcing the cluster to scale out or consuming a disproportionate number of containers in the cluster. Tenant 2, on the other hand, might be imposing minimal load.
The diagram below provides a conceptual model of how you might aggregate a series of calls to a service and use that data to determine a tenant's consumption. Here, we simply use call frequency to determine a percentage of activity for a tenant. This provides a model for distributing the compute costs associated with that service to each tenant based on those percentages.

This represents a highly simplified version of what you might build. Still, it gives a sense of what may be involved in arriving at a reasonable distribution of tenant compute costs. The good news here is that the mechanisms needed to support collecting this data overlap heavily with the general analytics tooling you'll want to put in place to analyze and assess tenant activity. The key is to make sure you've collected the necessary data with the context needed to support your cost allocation model.
Calculating storage costs
Although the previous example gave us a sense of how we might calculate compute consumption, that same model may not be a good fit for analyzing storage costs. A tenant might, for example, consume significant amounts of compute and still have minimal impact on storage costs. There is no guaranteed correlation between compute and storage consumption in SaaS environments.
Analyzing a metric like IOPS has parallels with what we discussed for analyzing compute consumption. The goal would be to derive some notion of a tenant's storage activity from that tenant's interactions with the data. These metrics can be derived from each tenant's interactions with a data access layer employed by your solution (as shown in the following diagram). With this approach, every call to acquire or manage data would be processed by a common framework that captures and aggregates tenant storage activity. That activity would then be used to apportion costs to each tenant.

Determining a tenant's impact on the system's storage footprint is a less dynamic process. Here, you may have data stored across several AWS services (Amazon DynamoDB, Amazon Simple Storage Service (Amazon S3), Amazon Redshift, Amazon Elastic Block Store (Amazon EBS) and so on), and you'll need to determine what percentage of that storage belongs to each tenant. Assessing this data typically requires a process that can periodically analyze the distribution of data to determine a tenant's storage consumption. The mechanism used for this analysis varies with your data partitioning model and the type of storage resource being evaluated. Amazon S3 and DynamoDB, for example, would require very different strategies for analyzing tenant storage consumption.
Billing metering vs. tenant consumption
Metering is an essential aspect of many SaaS environments. SaaS providers often build a billing and rating strategy around some dimension of consumption. These consumption dimensions cover a wide range of possibilities. Bandwidth, number of users, storage usage — all of these are kinds of billing models used to correlate tenant activity with some billing construct.
Metrics matter
Most SaaS businesses rely heavily on metrics. Having a pulse on your customers' often complex usage patterns is essential to understanding how to market, price, position and build your solution. Infrastructure overhead often represents a significant percentage of a SaaS provider's costs. So having an accurate view of how tenants are imposing load on your system gives both technical and business teams another variable that can shape the pricing and tiering strategies you choose to offer.
As you look at strategies for capturing and analyzing cost metrics, you should treat this as an iterative process. You can start with some very basic mechanisms to gain insight into tenant footprints that might otherwise go unnoticed. Then, over time, you can evolve and mature that model to add depth to your cost analysis.
The key is to bring visibility to the cost-per-tenant metric and make it part of the business's mental math.
In the next post, we'll talk about migration strategies for SaaS architectures.
SaaS Factory (The Series) — How to Build Software as a Service Solutions on AWS
See you in the next post! =)
Comments
Every comment is moderated before it appears here. Nothing is published automatically.
Loading…