Showing posts with label data management. Show all posts
Showing posts with label data management. Show all posts

Monday, February 24, 2014

ETL Technology in market

It is time to start tool based posts with Datastage and sharing knowledge. 

Before going deeper on Datastage, it is better to define what is ETL and what is current understanding and position of this technology in the market, for the ones new to the topic. 



ETL ("extract, transform, load") is used to define process or supporting tools that are used to pull data out of one source (database, file ...etc), make necessary transformations and load to another database as a target. As I stated in my previous post "Why data integration geting more important?" , there are main requirements that makes us to think on data integration. Mainly,
(-) Data is not a stand alone asset anymore for enterprises or organizations. 
(+) Data is a commodity moving around the enterprise and going in to and coming out from other processes, other systems, other enterprises...etc. 

To handle this commodity throughout your systems, enterprise and even your whole environment; instead of just worrying about full-compatibility between systems, you need to consider about your data integration capability. Compatibility and bundling of complementary tools is used as a marketing&sales strategy in the develeopment era of information technologies or mainly in computer science. However with the explosion of knowledge and technological develoments, it is almost impossible for a vendor to respond and meet all expectations in the market. Probably you already watch out that vendors more strictly stick to standards and more focused expertise together with partnership approach is being more popular. 

With the discusions on Bigdata and NoSQL systems, there are two main ideas on whether ETL will be still in use or not. You can find viewpoints of Phil Shelley, former CTO Sears Holdings, CEO Metascale who has also established his Bigdata consulting firm NPP-Newton Park Partners last year(2013) and James Markarian, CTO Informatica  from different sides in InformationWeek article.
It is impossible not to listen Shelley's idea that "since Hadoop came to the enterprise, we are beginning to see the end of ETL as we know it". But there are points that I do not completely agree with Shelley. While he is saying ETL, he is just focusing the technology as it is now but we know that technology is tend to evolve acording to requirements in the market and it is not necessarily to be called as a new technology. Additionally, I do not see each stage of ETL as  non-value-added activities. Within the context of relational databases and structured data, with a good design and good performance, you can add value to your data and turn your stand-alone asset to a commodity that can be used for different purposes throughout your organisation with ETL. Shelley might be completely right for Hadoop but I really suspect whether "Hadoop came to enterprise"! Although it comes to enterprise, is it possible to have Hadoop as the only system? He also states "Some subsets of data do have to be moved out of Hadoop into other systems, for specific purposes. However, with a strong and coherent enterprise data architecture, this can be managed to be the exception.

I do not want to make you lost in different articles before startig to learn a tool but I also strongly believe that it is better to understand the requirement and motivation behind any effort. ıt might also be good to have a look at "The State of ETL: Extract, Transform and Load Techology" article written by Alan R. Earls in DataInformed.

Not just for ETL but for all technologies, it is better to take it as a conceptual knowledge and make use of it to understand new technologies. It is more likely that ETL will not exist for too many years as its traditional form but the logic behind it to extract (filter and read, not necesarily to load to a different system) data, transform if necessary and to load (to a new system, to a modeling tool or just to a user interface) to make data available to serve specific purposes will retain. 

Hope you will find Datastage posts helpful for your ongoing tasks and to have a vision to get ready for new technologies. 

Tuesday, August 27, 2013

A quick look at Data Management

As enterprises are growing, data need at all levels and all types is increasing and decision making is getting more complicated. What business users see and ask for more is reports that they can access aggregated and analysed data in the form of information. On the other hand, what IT developers need to do to meet this expectation is much more than creating reports. 

Other than reporting purposes, data start to act as a commodity between different departments and processes in an enterprise and also between different enterprises. This extensive need for and use of data led the expansion of data management field and accelerated technological developments.


What we’re seeing now is an absolute explosion in data management technology and it’s come about because of the complexity of data problems 
Mike Olson, the CEO of Hadoop company Cloudera.

Data Management is defined as "Administrative process by which the required data is acquired, validated, stored, protected, and processed, and by which its accessibility, reliability, and timeliness is ensured to satisfy the needs of data users." in http://www.businessdictionary.com.

To understand the scope of data management and to get familiar with the field, we can have a look at  10 Data Management Functions defined by The Data Management Association.

  • Data Governance – The exercise of authority, control and shared decision-making (planning, monitoring and enforcement) over the management of data assets. Data Governance is high-level planning and control over data management.
  • Data Architecture Management – The development and maintenance of enterprise data architecture, within the context of all enterprise architecture, and its connection with the application system solutions and projects that implement enterprise architecture.
  • Data Development – The data-focused activities within the system development lifecycle (SDLC), including data modeling and data requirements analysis, design, implementation and maintenance of databases data-related solution components.
  • Database Operations Management – Planning, control and support for structured data assets across the data lifecycle, from creation and acquisition through archival and purge.
  • Data Security Management – Planning, implementation and control activities to ensure privacy and confidentiality and to prevent unauthorized and inappropriate data access, creation or change.
  • Reference & Master Data Management – Planning, implementation and control activities to ensure consistency of contextual data values with a “golden version” of these data values.
  • Data Warehousing & Business Intelligence Management – Planning, implementation and control processes to provide decision support data and support knowledge workers engaged in reporting, query and analysis.
  • Document & Content Management – Planning, implementation and control activities to store, protect and access data found within electronic files and physical records (including text, graphics, image, audio, video)
  • Meta Data Management – Planning, implementation and control activities to enable easy access to high quality, integrated meta data.
  • Data Quality Management – Planning, implementation and control activities that apply quality management techniques to measure, assess, improve and ensure the fitness of data for use.

In addition to defined functions, "Data Integration & Interoperability" is defined as "a new knowledge area in order to highlight its emerging importance in the field. (http://www.dama.org/i4a/pages/index.cfm?pageid=3733)

Data Management Knowledge Areas, DAMA
Data Management Knowledge Areas, DAMA




Related Posts Plugin for WordPress, Blogger...