Showing posts with label ER-modeling. Show all posts
Showing posts with label ER-modeling. Show all posts

Wednesday, March 20, 2013

SAP Powerdesigner: The beginning of the end? PRICE UPDATE

Introduction

In this post https://plus.google.com/u/0/106553443733559106071/posts/65vnSaUEioZ +Mirko Pawlak discusses a price hike of the sybase products like Powerdesigner (an ER modeling tool) of around 80%. Since Powerdesigner is already one of the most (if not THE most) expensive ER modeling tool this means Powerdesigner Licensing costs will go through the roof.

ER tools and standards

ER modeling tools and standards have always been an issue. Interoperability and standardization have never been very strong here, leading to expensive (but sometimes interesting) closed ecosystems and tooling applications. There are different notations and approaches like Barker and Information Engineering, some tools support several notations while others stick to one notation. ER tools have extended the notation to different aspects of information/data modeling and databases, so they could claim full support of the development cycle, even if this is not always what Chen had in mind when he developed ER modeling/diagramming. Transporting metadata from one tool to another has always been a costly endeavor, especially for not non basic metadata data of entities, attributes and keys. The only vendor seriously addressing this aspect is metaintegration: http://www.metaintegration.net/ which shows the how small and exclusive this market really is.

Powerdesigner vs the rest

Powerdesigner has always been one of my favorite tools, not because it was always the best, or that it did not have serious (sometimes even almost fatal) flaws, but because it is quite extensible and hence you can fix a lot of issues around diagramming and design, model generation and transformation. Other good tools are Embarcadero's ER/Studio, which lacks the integrated extensibility, but is a good tool with a reasonable automation interface. CA Erwin is one of the most active tool vendors out there with integrating and partnering. For example a Data Vault ER modeling notation and integration with BIReady. If this means Erwin is going to be customizable to the extent that powerdesigner has now I don't know.
Powerdesigner also has Enterprise Architecture, UML, XML, ETL, BPM and other notations under it's roof, but product integration with 3rd party tools (except database platforms) is not that extensive. However, integration into the SAP technology stack is going to be fixed in a short time frame, allowing Business Objects, SAP HANA and other tools to integrate into Powerdesigner.

Buying into Powerdesigner

Like Mirko mentioned, if you'd like to buy (more) powerdesigner (licenses) you might do it now (before the end of march). If not, plan B might just need a good dusting off. As with Business Objects, SAP customers are currently faced with a new (expensive) tool from outside the SAP realm with which they have no experience. The current (non SAP) customer base however needs to address the fact that their tool gets the most attention when integrating into the SAP ecosystem. I don't mind if this means that the foundation of powerdesigner becomes better, but I fear a lot of new developments will be specific extensions for product integration, that might give rise to more issues instead of less.

The future of Powerdesigner, ER modeling and database tools.

The market for data and database diagramming is constantly changing. I see a commodization of database development diagramming tools like Oracle SQL Developer or Microsoft SQL Server Data Tools. UML diagramming is also becoming a commodity. This alas does not imply that good data modeling is becoming a basic skill, since I see toolsets, vendors and notations ruling the data modeling space instead of good data modeling standards and foundations like the Relational Model and Fact Orientend Modeling (NIAM/FCO-IM). I also see more and more semantic tools and technologies as well, but connection to the database realm is opportunistic at best. In this light the expensive Case Tools and ER tools are not winning any war. But it does mean that one of the few tools out there that is flexible enough to accommodate all kinds of diagramming is becoming more of a niche tool in the SAP ecosystem.


PRICE UPDATE!

As shown at the toolpark website:  PD prices will go up indeed. As discussed here by +Mirko Pawlak  , it is just to match the prices of the competition (although licensing schemes and functionality are difficult to match) . For example, the old price of the basic Data Architect standalone seat (the cheapest license btw) goes from 2.440 Euro netto to 4.600 Euro netto.

It is clear to me that market forces here are driving the prices up instead of down. Which means I have to conclude that there is no real competitive market, as any specialist can attest to who has done a migration from one ER tool to another.





Monday, January 21, 2013

Data Vault and data modeling in the Trenches 27th,28th Feb.+ March 6

Intro

While I'm meeting a lot of Data Vault specialists in the field, there are always a lot of questions and assumptions unanswered while working in the trenches. The focus on value and productivity is difficult to merge with delving deep and doing reflection. The  DV in the trenches track is one of the few possibilities to dive into Data Vault without discussing urgent and pressing issues around loading and transforming your next Hub or Link. But of course, no theory survives contact with reality, so we will be doing a lot of reflection on actual implementations, current standards and best practices.

So if you really want to learn about data modeling and learn about Data Vault in a broader scope? Discuss Data Vault 2.0 or understand exotic but oh so practical key satellites? Getting to grips with 'modeling standards', architectures and frameworks? Anchor Modeling, or even what the rest of the MATTER program has to offer? Join me on the coming DV course of the DV in the trenches track.

Schedule

The MATTER program will be scheduling a new iteration of the Data Vault modeling course of the DV in the trenches track. It will be held February 27th and 28th, and an extra day on March the 6th 2013. I'll be teaching this course personally so I hope to see you there ;)

Registration

Registration can be done at the website of BI-Podium here.






Tuesday, November 20, 2012

Colors of The Data Vault

Introduction

Most Data Vault modelers color code their Data vault models. Alas, the chosen colors usually differ between (groups of) modelers. Hans Hultgren was one of the first to start color coding practice (see http://www.youtube.com/watch?v=kRoDRlj8_YU ), but others like Kasper de Graaf did also use these color coding as well.

The first practice to my knowledge was the following coding:
  • Hubs: Blue
  • Links: Red
  • Satellites: Yellow/Orange or Yellow/Green or Green
  • Reference Entities: None, but I use an unobtrusive grey/purple
I also use a shape coding as well:
  • Hubs: square/cube
  • Links: Oblong/oval
  • Sats: Rectangle/flat
  • Reference entity: circle
  • Hierarchical Link: pyramid
  • Same-As/From To link: Arrow like construct

Using Color Coding Practices

Color Coding can not only be done on the basic DV entities, but also on attributes, relationships and keys and hence on any style (Normalized/Dimensional) and other modeling technique (FOM's like FCO-IM). They provide an interesting visualization of modeling and transformation strategies. One of the best diagramming techniques to use color coding are FOM's like FCO-IM because it let's you track your transformation rules and you only need to color code roles and entities (whereas in ER you will be color coding keys, attributes and relationships)

Color Coding Data Models

Color coding data models is i fact a transformation strategy visualization method. It should rely on classification metadata of the underlying data model.
I'll show here some examples of color coding data models.

FCO-IM


In FCO-IM you color Nominalized fact types (hubs, links and cross reference tables) or all the roles of the fact type. You can aslo color code UC's role connectors as well

Data Vault

Here we usually just color code the basic entities.

DV Skeleton

The Dv skeleton is just the DV without the satellites.


3NF

In 3NF each and every entity can produce a hub, link and sat.


DV source system analysis

Here I show part of a detailed color coding of a 3NF source system that needs to be transformed to a Data Vault. See that not only entities, but also keys, attributes and relationships have extensive color coding as well as corresponding classification metadata.

Dimensional Model

A dimensional model shows a link as the basis for a fact table, with measures coming from a satellite. Dimensions are usually hubs with their accompanying satellites.

I hope this short overview will show you how you can use color coding the enhance the usefulness of your data (vault) model diagrams.

Thursday, October 11, 2012

Concerning Data Model Madness: what are we talking about?

There has been endless debates on how to name, identify, relate and transform the various kinds of Data Models we can, should, would, must have in the process of designing, developing and managing information systems (like Data warehouses). We talk about conceptual, logical and physical data models, usually in the context of certain tools, platforms, frameworks or methodologies. A confusion of tongues is usually the end result. Recently David Hay has made an interesting video (Kinds of Data Models -- And How to Name them) which he tries to resolve this issue in a consistent and complete manner. But on LinkedIn this was already questioned if this was a final or universal way of looking at such models.

One of the caveats is that a 'model' needs operators as well as data containers. Only the relational model (and some Fact oriented modeling techniques, which support conceptual queries) defines operators in a consistent manner. In this respect Entity Relationship techniques are formally diagramming techniques. Disregarding this distinction for now we need to ask ourselves if there are universal rules for kinds of Data Models and their place within a model generation/transformation strategy.

I note that if you use e.g. FCO-IM diagrams you can go directly from conceptual to "physical" DBMS schema if you want to, even special ones like Data Vault or Dimensional Modeling. I also want to remark that there are 'formal language' modeling techniques like Gellish that defy David's classification scheme. They are both ontological and fact driven and could in theory go from ontological to physical in one step without conceptual or logical layer (while still be consistent and complete btw, so no special assumptions except a transformation strategy). The question arises how many data model layers we want or need, and what each layer should solve for us. There is tension between minimizing the amount of layers while at the same time not overloading a layer/model with too much semantics and constructs which hampers its usability.

For me this is governed by the concept of concerns, pieces of interest that we try to describe with a specific kind of data model. These can be linguistic concerns like verbalization and translation, semantic concerns like identification, definition and ontology, Data quality constraints like uniqueness or implementation and optimization concerns like physical model transformation (e.g. Data Vault, Dimensional Modeling), partitioning and indexing. Modeling/diagramming techniques and methods usually have primary concerns that they want to solve in a consistent way, and secondary concerns they can model, but not deem important (but that are usually important concerns at another level/layer!). What makes this even more difficult is that within certain kinds of data models there is also the tension between notation, approach and theory (N.A.T. principle). E.g. the relational model is theoretically sound, but the formal notation of nested sets isn't visually helpful. ER diagramming looks good but there is little theoretic foundations beneath it.

I personally think we should try to rationalize the use of data model layers, driven by concerns, instead of standardizing on a basic 3 level approach of conceptual, logical, and physical. We should be explicit on the concerns we want to tackle in each layer instead of using generic suggestive layer names.

I would propose the following (data) model layers minimization rule:

A layered (data) modeling scenario supports the concept of separation of concerns (as defined by Dijkstra)  in a suitable way with a minimum of layers using a minimum of modeling methodologies and notations.