Komprise Reports: 3 Steps to Cost-Saving Insights

In our latest Office Hours blog, we cover Komprise Reports, with Benjamin Henry, Field CTO.  You can also read blogs on these sessions: intelligent data tiering, AI data workflows and KAPPA Data Services, unstructured data migration and Deep Analytics.

Register for upcoming sessions and watch the recordings on demand at komprise.com/office-hours.

The Pain: FinOps Isn’t a Separate Job Anymore

Storage cost accountability used to sit with one team. Now it belongs to whoever owns the data, which means storage, IT, and department leads all need to manage what Ben calls north-south-east-west, not just up to leadership. Most teams still build that case the hard way: pulling spreadsheets, downloading CSVs, chasing other departments for their numbers, every time someone asks what storage costs. That manual loop is exactly what Komprise Reports can remove.

Step One: Set the Cost Model Once

  • Build your cost model on the plan page. Hardware, backup, software, and number of backup copies retained turn into a per-gigabyte, per-month figure automatically, with separate settings for cloud versus on-premises targets.
  • Run more than one plan for more than one cost profile. Different vendors and different Opex and Capex structures require their own cost model and reports.
  • Don’t stop at the sticker price. Opex is easy to calculate per GB per month. Capex often hides real costs in year three and beyond of a longer depreciation cycle, plus rack space, power, cooling, and staff time for 24/7 care and drive swaps. Folding those in shows the true total cost of ownership, and a bigger savings number once cold data tiering is modeled against it.
  • Export as PDF for executives, CSV for everyone else. The CSV is a portable input for Snowflake, Databricks, Microsoft Fabric, or anything else that takes a spreadsheet.

Step Two: Show Back the Spend

  • ShowbackscreenshotBuild an executive showback report in a few clicks. Filter by site, share, or organization, or leave it blank to cover everything. It returns total data cost, a heat map, and top-consuming shares in a one-page format built for leadership, not just storage admins.
  • Save and schedule it. Once filtered the way you want, save the report, then schedule delivery on a recurring basis to any list of recipients: a site lead, a data steward, a research group.
  • Drill down to any individual share. The same report at the organization level and at a single team’s share level tells two different stories, one about total sprawl and the other about a specific team’s contribution to it.

Step Three: Turn the Report Into a Query, Then Into Action

  • Orphan data is one of the biggest avoidable costs. The orphaned data report filters files whose owner identity can no longer be resolved, typically because someone left the organization. Refine further by access history: some may still be in active use, so narrowing to data untouched for three years or more finds the safer data set for action.
  • Duplicate data search across every vendor. Most storage platforms only dedupe the data it stores. Komprise checks across the entire estate, with a configurable definition of duplicate: filename and extension, created time, or last modified time. Learn more about the potential duplicates report.
  • Every report view is also a Deep Analytics query. Reproduce the same filter as a saved query, and it can become the basis for a tiering, archive, or confine plan. A data steward can tighten “undefined owner” down to “not accessed in three years,” then apply a tag, with no ability to move or delete data directly.

The hard part usually isn’t finding the data. A finance organization managing petabytes with 96% cold data found that identifying the actual owners, not the data itself, was the real bottleneck, since storage teams alone couldn’t approve deletion or tiering. An orphaned-data report built on an undefined-owner query solves exactly that.

Governance Gets Its Own Reports Too

Give legal or compliance the ability to tag data subject to hold, so plans and workflows automatically skip it, and route it to its own storage zone that is covered by the legal or compliance department’s budget. Governance stakeholders can customize tags in any taxonomy  required such as “data sensitivity: ePHI,” “GDPR”. Users can add tags manually through a query, automatically through a PII or regex scan, or through the API using a third-party classification tool.

Scanning inside a file for sensitive content does not overwrite its access or modified timestamps, unlike some other vendors in this space, which matters because those same timestamps feed every other report in this post.

Why Now?

Flash pricing is at a premium and spinning-disk manufacturing capacity has been cut for years, so neither fallback is cheap right now. Buying another shelf is not always the responsible move when a chunk of what’s on the current shelf is ROT data doing nothing.

It’s elucidating to note that tagging and tiering, which powers a showback report, also sanitizes, classifies and curates data to the subset needed for AI. Query metadata can be extracted into Apache Iceberg tables through Transparent File Tables for Snowflake, Databricks, or Microsoft Fabric on a recurring schedule, not just a one-time CSV. Watch a demo.

Key Takeaways

  1. Set the cost model first, including the downstream Capex costs that are easily missed.
  2. Drill into Showback at the department and individual team share to get nuanced metrics.
  3. Orphan data and duplicate data reports convert directly into an actionable Deep Analytics query.
  4. Give end users scoped access, not full admin rights, so they can tag their own data and contribute to deeper accountability and savings.

Watch the Webinar

Read more Office Hours blog posts:

Watch all of the recordings and register for the next one at komprise.com/office-hours.

Getting Started with Komprise: